Display method and display system

By generating and controlling three-dimensional multimedia data through human-computer interaction, the problem of complex video switching in three-dimensional display is solved, the convenience and controllability of three-dimensional display are achieved, and the application scenarios and functions of three-dimensional display are expanded.

WO2025209102A1PCT designated stage Publication Date: 2025-10-09BOE TECHNOLOGY GROUP CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081046
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-02
Filing Date
2025-03-06
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

In a 3D display scenario, it is difficult for users to switch or transform the 3D video displayed on the large screen through human-computer interaction, which makes management and switching complicated.

Method used

Obtain interactive information through human-computer interaction operations, generate, play, switch, store, query and share multimedia data, including three-dimensional multimedia data, use voice, action and text input to generate virtual digital resources, combine resource generation models and three-dimensional synthesis models to realize content generation and display control of three-dimensional data.

Benefits of technology

It expands the role of human-computer interaction in 3D display, broadens the application scenarios and functions of 3D display, and enables users to generate and control 3D videos and images through simple human-computer interaction operations, providing a new and intuitive interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081046_09102025_PF_FP_ABST
    Figure CN2025081046_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A display method and a display system, relating to the technical field of data display. The display method comprises: acquiring interaction information, wherein the interaction information is generated on the basis of a human-computer interaction operation performed by a user; and executing a target operation on the bass of the interaction information, wherein the target operation comprises at least one of generating, playing, switching, storing, querying and sharing multimedia data, generating the multimedia data comprises acquiring on the basis of the interaction information a digital resource to be synthesized and generating three-dimensional-type multimedia data on the basis of the digital resource, and the digital resource comprises at least one of an audio resource, an image resource and a video resource.
Need to check novelty before this filing date? Find Prior Art

Description

Display method and display system

[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on April 2, 2024, with application number 202410396486.9 and invention name “Display Method and Display System,” the entire contents of which are incorporated by reference into this disclosure. Technical Field

[0002] The present disclosure relates to the field of display technology, and in particular to a display method and a display system. Background Art

[0003] Three-dimensional (3D) technology has shown great development potential in many fields, such as consumer electronics, medical imaging, etc. Currently, a variety of 3D display technologies have emerged, such as 3D display technologies using glasses and 3D display technologies without glasses.

[0004] Overview

[0005] Based on the background technology, the present disclosure proposes a display method and a display system.

[0006] The present disclosure provides a display method, the method comprising:

[0007] Acquiring interaction information, where the interaction information is generated based on a human-computer interaction operation performed by a user;

[0008] performing a target operation based on the interaction information, the target operation comprising at least one of generating, playing, switching, storing, querying, and sharing multimedia data;

[0009] Generating multimedia data includes: obtaining digital resources to be synthesized based on the interactive information, and generating three-dimensional multimedia data based on the digital resources; the digital resources include at least one of audio resources, image resources and video resources.

[0010] Exemplarily, the human-computer interaction operation includes at least one of a voice input operation, a body gesture operation, and a text input operation.

[0011] Exemplarily, the acquiring the digital resource to be synthesized based on the interaction information includes:

[0012] Parsing the interaction information to obtain first text information, where the first text information is used to indicate a location and / or identifier of a digital resource;

[0013] The digital resource is acquired based on the first text information.

[0014] Exemplarily, the interaction information is generated according to an opening operation of a multimedia resource on a terminal, and acquiring the digital resource based on the first text information includes:

[0015] Determining a first data source where the data resource is located based on the first text information;

[0016] Acquire a first digital resource from the first data source, and acquire a second digital resource from a second data source different from the first data source; wherein a similarity between the second digital resource and the first digital resource is greater than a preset similarity;

[0017] The first digital resource and / or the second digital resource are used as the digital resources to be synthesized.

[0018] Exemplarily, the acquiring the digital resource to be synthesized based on the interaction information includes:

[0019] Parsing the interaction information to obtain text information, where the text information is used to at least indicate content of the multimedia data;

[0020] Inputting the text information into at least one resource generation model to obtain a virtual digital resource generated by the resource generation model; wherein the resource generation model is trained using a plurality of text samples and virtual digital resource samples corresponding to each text sample as training samples;

[0021] The virtual data resources include audio type resources and / or image type resources.

[0022] Exemplarily, the human-computer interaction operation is a voice input operation, and parsing the interaction information to obtain text information includes:

[0023] Inputting the interactive information into a content generation model to obtain the text information output by the content generation model;

[0024] The content generation model is obtained by training using a plurality of interaction information samples as training samples.

[0025] Exemplarily, generating the multimedia data based on the digital resource includes:

[0026] Acquire an image resource belonging to an image type from the digital resources, and acquire an audio resource belonging to an audio type from the digital resources;

[0027] Generate video resources based on the image resources;

[0028] The video resource and the audio resource are synthesized to obtain the multimedia data.

[0029] Exemplarily, generating a video resource based on the image resource includes:

[0030] Inputting the image resource into a three-dimensional synthesis model to obtain at least one video resource output by the three-dimensional synthesis model;

[0031] The three-dimensional synthesis model is obtained by training with a variety of two-dimensional image samples as training samples and three-dimensional multimedia data corresponding to the two-dimensional image samples as supervision.

[0032] Exemplarily, the interactive information includes at least one style of multimedia data; and inputting the image resource into the three-dimensional synthesis model includes:

[0033] Inputting the image resource and the style in the interaction information into the three-dimensional synthesis model to obtain video resources corresponding to each style;

[0034] And / or, the image resource is input into a three-dimensional synthesis model corresponding to each of the styles to obtain a video resource output by each of the three-dimensional synthesis models, wherein different three-dimensional synthesis models correspond to different styles.

[0035] Exemplarily, the acquiring interaction information includes:

[0036] receiving interaction information sent by a target program located on a first terminal, wherein the target program is configured to send the interaction information in response to a human-computer interaction operation performed by a user, and the interaction information carries a user account on the first terminal;

[0037] The playing of the multimedia data includes:

[0038] Based on the user account, determining a second terminal bound to the user account;

[0039] The multimedia data is sent to the second terminal so as to play the multimedia data on the second terminal; wherein the second terminal is the first terminal or a terminal different from the first terminal.

[0040] Exemplarily, the interaction information is sent by the third terminal after identifying the user's action in response to the operation of opening the multimedia resource on the third terminal; wherein the user's action includes at least one of a gesture action, a body action, and a head action.

[0041] Exemplarily, the interaction information carries the first account and the second account, and the sharing of the multimedia data includes:

[0042] displaying a variety of multimedia data associated with the first account;

[0043] In response to a selection operation on the associated multiple multimedia data, the selected first multimedia data is sent to the terminal where the second account is located.

[0044] Exemplarily, the interaction information carries an identifier of the second multimedia data to be converted, and the switching of the multimedia data includes:

[0045] Display multiple preset styles;

[0046] In response to a selection operation of the plurality of preset styles, the second multimedia data is converted into third multimedia data conforming to a target style; wherein the target style is the selected preset style.

[0047] Exemplarily, the interaction information carries a fourth terminal to be displayed, and the switching of the multimedia data includes:

[0048] Acquiring fourth multimedia data currently played by the terminal that triggers the interaction information;

[0049] Acquire a display type of a display screen configured for the fourth terminal, where the current display type includes a three-dimensional display type and a two-dimensional display type;

[0050] Based on the display type and the type of the fourth multimedia data, the fourth multimedia data is displayed on the fourth terminal.

[0051] The present disclosure further provides a display system, comprising a server and a terminal connected to the server, wherein the terminal is configured with a three-dimensional display screen, wherein:

[0052] The server is configured to obtain interaction information and perform a target operation on the multimedia data based on the interaction information; wherein the target operation includes at least one of generating, playing, switching, storing, querying, and sharing the multimedia data; the interaction information is generated based on at least one of a voice input, an action performed, and a text input by the user; and the digital resource includes at least one of an image resource, an audio resource, and a video resource;

[0053] The terminal is used to play the multimedia data after performing the target operation.

[0054] Exemplarily, the system is further configured with a target program, which is set on the terminal or on an electronic device connected to the terminal; wherein, when the target program is run, it is used to implement the following steps:

[0055] Sending the interaction information to the server based on the target operation input by the user;

[0056] The target operation includes at least one of a voice operation, a gesture operation, and a text input operation.

[0057] Exemplarily, the server is configured with a resource generation model, and the resource generation model is configured to generate virtual digital resources based on input text information; wherein the virtual data resources include audio type resources and / or image type resources;

[0058] The resource generation model is obtained by training using a variety of text samples and virtual digital resource samples corresponding to each text sample as training samples.

[0059] Exemplarily, the server is configured with at least one 3D synthesis model, wherein the 3D synthesis model is used to generate a 3D image resource based on an input image resource;

[0060] Among them, different three-dimensional synthetic models correspond to different styles.

[0061] Exemplarily, the terminal is configured with a motion recognition module, and the motion recognition module is configured to, in response to an operation of opening a multimedia resource on the third terminal, recognize a user's motion and then send the interaction information;

[0062] The action includes at least one of a gesture action, a body action, and a head action.

[0063] Exemplarily, the server is further configured with a database storing a variety of digital resources; the terminal is further configured with a resource library including a variety of multimedia resources;

[0064] The digital resources are derived from the resource library and / or the database.

[0065] The display method provided by the present disclosure can obtain interactive information; based on the interactive information, perform target operations, wherein the target operations include generating, playing, switching, storing, querying, and sharing multimedia data; wherein generating multimedia data includes: obtaining digital resources to be synthesized based on the interactive information, and generating three-dimensional multimedia data based on the digital resources; the digital resources include at least one of image resources, audio resources, and video resources, and the interactive information is generated based on at least one of user input voice, executed actions, and input text. Since at least one of the operations of generating, playing, switching, storing, querying, and sharing multimedia data can be performed directly based on the user's interactive operation during display, for example, at least three-dimensional multimedia data can be generated, thus, in 3D display technology, the indicated three-dimensional video and / or three-dimensional image can be automatically synthesized based on the acquired interactive operation, thereby integrating human-computer interaction into the display control and content generation of the three-dimensional display, thereby providing users with a new, intuitive, and efficient interactive experience.

[0066] The above description is only an overview of the technical solution of the present disclosure. In order to more clearly understand the technical means of the present disclosure, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present disclosure more obvious and easy to understand, the specific implementation methods of the present disclosure are listed below.

[0067] BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following is a brief introduction to the drawings required for the description of the embodiments or related technologies. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. It should be noted that the scales in the drawings are for illustration only and do not represent the actual scale.

[0069] FIG1 shows a schematic diagram of a communication scenario in an embodiment of the present disclosure;

[0070] FIG2 is a schematic diagram showing a flow chart of steps of a display method according to an embodiment of the present disclosure;

[0071] FIG3 shows the types of human-computer interaction operations supported by a terminal in an embodiment of the present disclosure;

[0072] FIG4 shows a schematic diagram of an implementation method for acquiring digital resources in an embodiment of the present disclosure;

[0073] FIG5 shows a schematic diagram of an implementation method for generating multimedia data according to an embodiment of the present disclosure;

[0074] FIG6 shows a schematic diagram of a display scene in an embodiment of the present disclosure;

[0075] FIG7 shows another schematic diagram of a display scene in an embodiment of the present disclosure;

[0076] FIG8 shows a schematic diagram of another display scene in an embodiment of the present disclosure;

[0077] FIG9 shows a flow chart of a display method according to Example 1 of an embodiment of the present disclosure;

[0078] FIG10 is a schematic diagram showing a flow chart of a display method according to Example 2 in an embodiment of the present disclosure;

[0079] FIG11 shows a schematic diagram of the layout structure of a display system according to an embodiment of the present disclosure;

[0080] FIG12 shows a structural diagram of an exemplary display system according to an embodiment of the present disclosure.

[0081] Detailed description

[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0083] In related technologies, the rapid development of three-dimensional display technology has built a bridge for the interaction between the real world and the virtual world. Products and systems based on three-dimensional display technology are also constantly being updated. However, in three-dimensional display scenarios, most research focuses on how to satisfy users' needs to manipulate virtual objects based on human-computer interaction while viewing them. In practice, the application scenarios of three-dimensional display technology are diverse. For example, in the field of consumer electronics, three-dimensional videos need to be displayed on large three-dimensional display screens set up in shopping malls. For example, in some art museums, large three-dimensional display screens can be set up in the museum to display artistic three-dimensional videos or images. In the above scenarios, the three-dimensional videos displayed on the large screen are generally fixed; therefore, it is difficult for users to switch, transform, and perform other operations on the three-dimensional videos displayed on the large screen through human-computer interaction, which makes the management and switching of the three-dimensional videos displayed on the large screen more complicated.

[0084] In view of this, the present disclosure proposes a method for applying human-computer interaction operations to the display control, content generation, switching and control of three-dimensional display, so that the system can perform content generation, content switching and display control of three-dimensional data based on human-computer interaction operations, thereby expanding the role of human-computer interaction in three-dimensional display and expanding the application scenarios and functions of three-dimensional display.

[0085] First, some technical terms used in the following embodiments of the present disclosure are explained:

[0086] Human-Machine Interaction (HMI): This is a subfield of computer science that studies how humans interact with computers and other machines, and how to design user-friendly interfaces to enable such interactions. In this patent, HMI encompasses how to interact with 3D display devices using voice input and gesture control.

[0087] Autostereoscopic Display: This is a display technology that allows 3D effects to be seen directly with the naked eye, without the need for 3D glasses or other auxiliary equipment. This technology is usually achieved by creating a parallax effect on the display, so that the left and right eyes see slightly different images, thus creating a 3D visual effect.

[0088] Speech-to-Text Model (STT): Also known as a speech recognition model, this is an AI model that takes speech input and converts it into text. This model is commonly used in voice-assisted technologies, such as smart assistants and automatic speech transcription services.

[0089] GPT-like text models: GPT (Generative Pretrained Transformer) is a natural language processing model developed by OpenAI. It has two stages, pre-training and fine-tuning, and can be used to generate human language. GPT-like text models are based on the GPT architecture, but may have been modified and optimized to suit specific application needs.

[0090] Diffusion Image Generation Model: This is an artificial intelligence image generation technology that generates new images by simulating the gradient diffusion process (that is, starting from random noise and gradually forming the final target image).

[0091] SoC (System-on-a-Chip): An SoC is an integrated circuit (IC) that integrates most or all required hardware functions, such as a central processing unit (CPU), memory, and input / output (I / O) interfaces, on a single chip. This design approach can reduce manufacturing costs and power consumption while improving performance.

[0092] Digital Asset: Anything that can exist in digital form and has value can be considered a digital asset. In this patent, digital assets may refer to resources such as 3D models, audio files, and text data that can be displayed and interacted with on a naked-eye 3D display device.

[0093] Gesture interaction: This is a human-computer interaction method that analyzes the user's body movements, especially hand and finger movements, and converts them into commands that the machine can understand. Gesture recognition technology can be used in many different application scenarios, including virtual reality, augmented reality, computer games, medical diagnosis, and robotics.

[0094] Voice interaction: This is another form of human-computer interaction. It recognizes and interprets user voice commands, translating them into commands that the machine can understand. Voice interaction technology is a fundamental feature of many modern devices (such as smartphones, smart speakers, and in-car infotainment systems), allowing users to operate with gestures and without visual constraints.

[0095] Light Field Display: Light Field Display is an advanced 3D display technology that creates a 3D effect by simulating light rays emitted from every viewpoint (called a "light field"). Compared to traditional 3D display technologies (such as stereoscopic vision), Light Field Display not only produces a realistic 3D depth perception, but also does not require the user to wear special 3D glasses. Therefore, Light Field Display has the potential to become the mainstream technology for naked-eye 3D displays in the future.

[0096] OLED (Organic Light Emitting Diode Display): OLED is a light-emitting diode technology that uses organic materials as the light-emitting layer. Compared to traditional liquid crystal displays (LCDs), OLED screens offer more vivid colors, deeper blacks, and wider viewing angles. Furthermore, because OLED pixels can emit light independently, OLED screens can be made thinner and lighter.

[0097] Infrared Camera: An infrared camera is an image device that can capture infrared light. It is commonly used in night vision, thermal imaging, and certain biological applications.

[0098] 1 and 2 , FIG1 shows a schematic diagram of a communication scenario of the display method of the present disclosure, and FIG2 shows a schematic diagram of the step flow of the display method of the present disclosure. As shown in FIG1 and FIG2 , the display method of the present disclosure can be specifically applied to a server. Specifically, the server can be a server or a cloud composed of a server cluster, which can include the following steps:

[0099] Step S101: Acquire interaction information; wherein the interaction information is generated based on a human-computer interaction operation performed by a user;

[0100] Step S102: performing a target operation based on the interaction information; wherein generating multimedia data includes: acquiring a digital resource to be synthesized based on the interaction information, and generating three-dimensional multimedia data based on the digital resource;

[0101] The target operation includes at least one of generating, playing, switching, storing, querying and sharing multimedia data;

[0102] The digital resources include at least one of audio resources, image resources and video resources.

[0103] In this embodiment, human-computer interaction operations may include voice interaction operations, action interaction operations, and text interaction operations, and the interaction information may be generated based on at least one of the user's voice input, the action performed, and the input text. The content and type of the interaction information may vary depending on the human-computer interaction operation.

[0104] Referring to FIG3 , a type of human-computer interaction operation supported by a terminal is shown. As shown in FIG3 , the form of human-computer interaction is not limited to the forms provided above, and other forms are also possible. For example, it can be brain-computer human-computer interaction, such as a user can wear a brain-computer device, which can collect EEG signals and process the EEG signals into data signals and then feed them back to the server, or directly feed the EEG signals as interaction information to the server. Therefore, the interaction information can also be EEG signals.

[0105] For example, human-computer interaction operations may include voice input operations, motion operations, and text input operations. For different operations, the type and content of interaction information may vary slightly. For example, if the interaction operation is a voice input operation, the server may obtain the voice uploaded by the user. In this case, the interaction information may be audio information. For another example, if the human-computer interaction operation is a motion operation, such as a user's gesture operation, the interaction information obtained by the server may be image information uploaded by an image acquisition device based on the user's gesture operation. For another example, if the human-computer interaction operation is a text input operation, the interaction information obtained by the server may be text information.

[0106] Of course, in some other examples, human-computer interaction operations can also be control operations performed by users with the help of control devices, such as interactive operations performed with a remote control. In this case, the interaction information can also be other types of information. For example, if the user uses a remote control for human-computer interaction, the interaction information obtained by the server can also be instruction information, such as instruction information issued by the remote control.

[0107] In this embodiment, different interaction information can be used to indicate different target operations on multimedia data, where the target operations include at least one of generating, playing, switching, storing, querying, and sharing multimedia data. Specifically, one type of interaction information can indicate the performance of at least one of the aforementioned operations. In this way, one or more operations on multimedia data can be implemented through a single human-computer interaction.

[0108] Multimedia data can include image data, video data, and audio data. Image data is not limited to real-life images, artistic images, or device-generated images, and can be two-dimensional or three-dimensional. Video data is not limited to three-dimensional or two-dimensional video, and is not limited to the type or source of the video data. Audio data can be synthesized audio data or natural human speech.

[0109] Specifically, the multimedia data may be located on a server, a terminal, or a third-party platform different from the server and the terminal, and no special restrictions are imposed here.

[0110] In the case of playing multimedia data based on interactive information, the multimedia data to be played can be determined based on the interactive information, then retrieved from the device where the multimedia data to be played is located, and then played on the designated device.

[0111] The switching of multimedia data may include: switching one type of multimedia data to another type of multimedia data, switching the style of multimedia data to another style, and switching the terminal that plays the multimedia data to another terminal. When switching multimedia data based on interaction information, the multimedia data before and after the switching and the terminals before and after the switching may be determined, so that the switched multimedia data is played on the terminal, or the multimedia data is switched to another terminal for playback.

[0112] Storing multimedia data may refer to storing the multimedia data specified by the user in a specified address, such as storing it in the address of a specified terminal, storing it in a specified account, and so on.

[0113] Querying multimedia data may refer to querying the multimedia data from multiple data sources according to query keywords carried in the interactive information, and feeding back the query multimedia data to the user.

[0114] Sharing multimedia data can include sharing multimedia data on a user account with multiple other users, or sharing multimedia data on a terminal with multiple other terminals. In this case, the interaction information can include the multimedia data to be shared, as well as the user or terminal to which the multimedia data is to be shared, thereby enabling multimedia data to be shared between users and between terminals.

[0115] The multimedia data targeted by the switching, storage, query and sharing can be three-dimensional multimedia data, such as three-dimensional video and three-dimensional image, or two-dimensional multimedia data, such as two-dimensional video and two-dimensional image.

[0116] In this embodiment, when generating multimedia data, the user may first obtain the digital resources required for generating the multimedia data based on the interactive information. The digital resources may be understood as the material for generating the multimedia data. The generated multimedia data may be 3D multimedia data, such as 3D video or 3D image.

[0117] In this example, the interactive information can carry content descriptions and style descriptions of the three-dimensional multimedia data. The content description can be used to describe the image and video content included in the media data, such as descriptions of characters, animal types, and environments in the video. In some examples, it can also include a description of the video's plot. The digital resources obtained based on the interactive information can be resources that match the content and style descriptions included in the interactive information. These digital resources can be used to synthesize multimedia data containing the content and style descriptions.

[0118] Among them, digital resources may include at least one of audio resources, image resources and video resources. Specifically, it may only include audio resources. When audio resources are included, the multimedia data generated may be audio data with changed frequency to create a voice-changing effect; when image resources are included, the multimedia data generated may be data after three-dimensionalization of the image; when video resources are included, the multimedia data generated may be three-dimensional data.

[0119] Specifically, the image resources included in the digital resources can be two-dimensional image resources, and the video resources can be two-dimensional videos. Therefore, when generating three-dimensional multimedia data based on digital resources, the two-dimensional digital resources do not need to occupy more storage resources, and based on the user's human-computer interaction operations, 3D videos and 3D images are generated.

[0120] Among them, image resources may include multiple images, audio resources may also include multiple audios, and video resources may also include multiple videos. It should be noted that since the acquired digital resources are obtained based on interactive information, there is a certain correlation between the multiple images, audio resources and video resources. The correlation may refer to the correlation between the content included in the image, such as including the same content; the correlation between audios in frequency and content, such as multiple audios belonging to the same frequency band or targeting the same content; the videos have correlation in content and plot, such as including similar content and similar plots.

[0121] Among them, after the multimedia data is synthesized, the multimedia data can be played on the terminal. As a result, the user only needs to perform human-computer interaction operations, and the server can automatically obtain digital resources and synthesize multimedia data without the user performing complex operations, thereby improving the convenience and controllability of 3D display.

[0122] By adopting the display method of this embodiment, at least one of the operations of generating, playing, switching, storing, querying and sharing multimedia data can be performed based on the human-computer interaction operation performed by the user, and the multimedia data can be three-dimensional multimedia data. Therefore, the human-computer interaction operation can be applied to the display control, content generation, switching and control of three-dimensional display, so that the system can perform content generation, content switching and display control of three-dimensional data based on the human-computer interaction operation, thereby expanding the role of human-computer interaction in three-dimensional display and expanding the application scenarios and functions of three-dimensional display.

[0123] In some embodiments, the human-computer interaction operation may include at least one of a voice input operation, a body gesture operation, a text input operation, and a brain-computer interaction operation. An interaction message may be generated based on a single human-computer interaction operation, or may be generated based on multiple human-computer interaction operations.

[0124] For example, if the user performs a voice input operation and a text input operation successively, where the voice input operation can be used to record a segment of audio, and the text input operation is used to indicate the processing of the recorded audio, such as what operation to perform based on the recorded audio to obtain multimedia data, then the interactive information can carry the audio and the processing of the audio.

[0125] As another example, if the user performs a gesture operation and a text input operation successively, where the gesture operation is used to indicate the digital resource to be processed and the text input operation indicates the content and style of the multimedia data, then the interaction information may include the identifier of the digital resource to be processed, the content and style of the multimedia data. In this way, the digital resource can be processed into three-dimensional multimedia data that conforms to the corresponding content and style based on the interaction information.

[0126] In combination with the above embodiments, the brain-computer interaction operation can be an interaction completed by the brain-computer device by collecting the user's electroencephalogram signals when the user wears the brain-computer device.

[0127] In some other embodiments, the digital resources may be obtained by acquiring from a data source, automatically generating, etc. When the digital resources are automatically generated, virtual digital resources may be generated based on the interactive information.

[0128] Specifically, in one example, digital resources can be directly pulled from the data source. Specifically, the interactive information can be parsed to obtain first text information, which is used to indicate the location and / or identification of the digital resource; based on the first text information, the digital resource is obtained.

[0129] The first text message may include the location of the digital resource, which may be the storage address of the digital resource, such as a URL, or the sector location of the digital resource on the terminal. Alternatively, the first text message may include an identifier of the digital resource, which may uniquely identify the digital resource and, through the identifier, may be used to locate the digital resource in a massive database.

[0130] When obtaining a digital resource based on the first text information, the digital resource can be directly called from the storage device where the digital resource is located, or the digital resource can be found from a massive database based on an identifier. Alternatively, if the first text information includes an identifier and a location, the digital resource can be obtained based on both. In this case, the location of the digital resource can be the location of the storage area where the digital resource is located, such as the address of a large storage area. Then, based on the identifier, the required digital resource can be found from multiple digital resources stored in the storage area. This method can narrow the search scope and improve query efficiency when the specific address of the digital resource is not clear.

[0131] In a further example of this embodiment, digital resources can be pulled out from multiple data sources. Specifically, after finding the digital resources indicated by the first text information, in order to expand the digital resources and enhance the diversity of data resources, the first digital resource can be obtained online, and then multiple digital resources (second digital resources) associated with or similar to the first digital resource can be obtained to enhance the diversity of digital resources.

[0132] Specifically, based on the first text information, the first data source where the digital resource is located can be determined; and the first digital resource can be obtained from the first data source, and the second digital resource can be obtained from a second data source that is different from the data source; then, the first digital resource and / or the second digital resource can be used as the digital resource to be synthesized.

[0133] The similarity between the second digital resource and the first digital resource is greater than a preset similarity.

[0134] In this example, the first data source may be different from the second data source, or the second data may originate from the same source as the first data source. Here, no special limitation is imposed on the relationship between the first data source and the second data source.

[0135] The first text message may include the first data source where the digital resource resides. Thus, during human-computer interaction, the user can specify which data source to obtain the digital resource from, that is, specify the first data source, and then specify which digital resources from the first data source to obtain. Thus, after interpreting the first text message, the first data source where the digital resource resides can be obtained, and the first digital resource can then be obtained from the first data source.

[0136] After obtaining the first digital resource, in order to expand the diversity of digital resources, a similar second digital resource can be obtained from other data sources or from the first data source. The second data source can be located on the server side or on the terminal side, and there is no limitation here. The similarity can refer to the degree of similarity between the content of the second digital resource and the content of the first digital resource. For example, if the first digital resource is an image resource, the similarity can refer to the similarity between images, and the similarity can be represented by cosine distance. For another example, if the first digital resource is an audio resource, the similarity can be the similarity between audios, such as the similarity between speech frequencies and speech contents; for another example, if the first digital resource is a video resource, the similarity can refer to the similarity between video images.

[0137] The digital resource with a similarity greater than a preset similarity may be used as the second digital resource.

[0138] When the second digital resource is obtained, the first digital resource and the second digital resource can be used as digital resources to be synthesized to participate in the synthesis of multimedia data; or, only the second digital resource or the first digital resource can be used as the digital resource to be synthesized, and when the second digital resource is used as the digital resource to be synthesized, the second digital resource can come from a second data source that is different from the first data source. Therefore, when the user specifies a digital resource on a certain data source, other similar but completely identical digital resources can be given, which can help the user expand the diversity of multimedia data, so that the user can obtain other similar and similar digital resources based on one digital resource.

[0139] In another example, when performing human-computer interaction operations, the user does not need to specify the digital resources or the data source where the digital resources are located. Instead, the user can specify the content of the desired digital resources, and the cloud will automatically generate the digital resources based on the indicated content.

[0140] In this example, when acquiring digital resources based on interactive information, the interactive information can be parsed to obtain text information; the text information is input into at least one resource generation model to obtain a virtual digital resource generated by the resource generation model; wherein the resource generation model is trained using multiple text samples and virtual digital resource samples corresponding to each text sample as training samples;

[0141] The text information is used to at least indicate the content and style of the multimedia data, and the virtual data resources include audio type resources and / or image type resources.

[0142] Referring to Figure 4, a schematic diagram of an implementation method for obtaining digital resources is shown. As shown in Figure 4, in this embodiment, the interactive information can carry the content and style included in the required digital resources. As mentioned above, depending on the type of digital resources, the content can be picture content, audio content, video content, etc., and the video content can include picture content and the plot of picture changes, etc.; among them, depending on the type of digital resources, the style can include the color style of the picture, the painting style, the background of the times, etc.

[0143] In this example, the interactive information can be text or voice, and can be generated based on the text and voice input of the user. For example, the user can enter the text "Generate a cute animation of a pig" or the voice input "Generate an animated video of a flying pig." Regardless of the type of interactive information, the interactive information can be parsed to obtain text information. For example, if the interactive information is voice, the interactive information can be subjected to voice recognition to obtain text information. If it is text information, the text information can be segmented to obtain text information containing keywords.

[0144] Specifically, the text information may include keywords indicating the content of the multimedia data, as well as keywords indicating the style thereof. Through a variety of keywords, the content and style of the digital resource required by the user may be acquired.

[0145] After obtaining the text information, the text input can be input into a resource generation model to obtain a virtual digital resource. The resource generation model can be used to generate an image-type digital resource or an audio-type digital resource. For example, different types of digital resources can be generated by different resource generation models. For example, the resource generation model can include a model for generating a virtual digital resource for audio and a model for generating a virtual digital resource for images.

[0146] The text information can specify the type of digital resource to be generated. Specifically, a text message can specify the generation of both audio and image virtual digital resources, or it can specify the generation of either audio or image virtual digital resources. When it is necessary to generate both audio and image virtual digital resources, the text information can be input into both resource generation models to generate audio and image virtual digital resources, respectively.

[0147] When using the resource generation model to generate virtual digital resources, one virtual digital resource or multiple virtual resources can be generated. Specifically, when the text information specifies that the multimedia data is video data, multiple image-type virtual digital resources need to be generated to synthesize the video. When the text information specifies that the multimedia data is three-dimensional data, multiple image-type virtual digital resources also need to be generated. These multiple virtual digital resources can be digital resources from different perspectives to synthesize the three-dimensional multimedia data.

[0148] The resource generation model is trained using multiple text samples and virtual digital resource samples corresponding to each text sample as training samples. For example, for a model generating image-type virtual digital resources, its virtual digital resource samples may be image samples, while for a model generating audio-type virtual digital resources, its virtual digital resource samples may be audio samples.

[0149] For example, a user enters the text "Generate a cute animation of a pig." After parsing, the text information "pig, cute, animation" is obtained. This information is input into the resource generation model, resulting in multiple images of pigs. These images can be virtual images, that is, synthesized animation images. Thus, the generated multiple virtual digital resources can be used to synthesize a video animation of the pig.

[0150] Among them, when training the resource generation model, the text sample can be input into the resource generation model to obtain the predicted digital resource output. Then, the virtual digital resource sample corresponding to the input text sample is used as the label, and the difference between the predicted digital resource and the virtual digital resource sample is calculated. The difference is used as the loss, and the resource generation model is updated based on the loss. After multiple updates, the resource generation model can be obtained.

[0151] In some examples, the analysis of interaction information can also be performed by a model. Specifically, the interaction information can be input into a content generation model to obtain text information output by the content generation model, wherein the content generation model is trained using multiple interaction information samples as training samples.

[0152] In one implementation of this example, the content generation model can be an STT model. Since the interaction information is generated based on voice input, it can be voice information. Therefore, the voice-type interaction information can be directly uploaded to the cloud, where the STT model parses the voice information to obtain text information. In this approach, the interaction information samples can be voice samples, and during the training of the content generation model, the text information samples corresponding to the interaction information samples can be used as supervision to construct a loss function, thereby calculating the loss value and updating the model parameters.

[0153] In another implementation of this example, the content generation model can be a feature extraction model that can be used to extract keywords from interaction information to obtain text information. In this approach, the interaction information samples can be text samples. During the training of the content generation model, the text information samples corresponding to the interaction information samples can be used as supervision to construct a loss function, thereby calculating the loss value to update the model parameters.

[0154] In one embodiment, the synthesized multimedia data is three-dimensional data, which may include image-type three-dimensional data or video-type three-dimensional data. For image-type three-dimensional data, the digital resources may include both image resources and audio resources. The process of obtaining the image resources and audio resources may refer to the description of the above example. When generating the multimedia data, an image resource of the image type may be obtained from the digital resources, and an audio resource of the audio type may be obtained from the digital resources. A video resource may be generated based on the image resource. Finally, the video resource and the audio resource may be synthesized to obtain the multimedia data.

[0155] In this embodiment, the synthesized multimedia data is a three-dimensional video, and its synthesis requires image resources and audio resources. Specifically, the image resources can refer to images, and the multimedia data includes multiple frames of three-dimensional images. Each frame of the three-dimensional image is synthesized by images from multiple different perspectives. Then, the synthesized video resources and audio resources are fused to obtain multimedia data. The fusion of its video and audio can refer to relevant technologies and will not be elaborated here.

[0156] Therefore, the acquired digital resources can be used to synthesize three-dimensional video data, thereby achieving the automatic generation of multimedia data from scratch. As a result, users can obtain three-dimensional videos that did not originally exist in the terminal and server through human-computer interaction operations. This makes it possible to integrate human-computer interaction into the content generation of three-dimensional videos, thereby providing users with a new generative experience.

[0157] In a further example of this example, when generating video resources based on image resources, they can also be generated according to a neural network model. Specifically, the image resources can be input into a three-dimensional synthesis model to obtain at least one video resource output by the three-dimensional synthesis model; wherein the three-dimensional synthesis model is trained using a variety of two-dimensional image samples as training samples and the three-dimensional multimedia data corresponding to the two-dimensional image samples as supervision.

[0158] Referring to Figure 5, a schematic diagram of an implementation method for generating multimedia data is shown. As shown in Figure 5, the three-dimensional synthesis model can be a 2D to 3D model, which can be used to convert a two-dimensional image into a three-dimensional image. The three-dimensional synthesis model can be used to convert input two-dimensional images of multiple different perspectives for a video frame into a three-dimensional image, thereby obtaining multiple frames of three-dimensional images, and then converting the multiple frames of three-dimensional images into a three-dimensional video resource.

[0159] The 3D synthesis model can be deployed in the cloud and trained using a variety of 2D image samples as training examples and the corresponding 3D multimedia data as supervision. During training, the 2D image samples are input into the 3D synthesis model to produce synthesized 3D images. A loss is then calculated based on the 3D multimedia data and the 3D images, and the parameters of the 3D synthesis model are updated based on the loss.

[0160] In some implementations of this example, the interaction information may indicate the style of the multimedia data. When synthesizing the multimedia data, the digital assets may be synthesized into multimedia data that conforms to the indicated style. As described above, the style may include the color style, painting style, and historical context of the multimedia data. Thus, when the 3D synthesis model generates video assets, the corresponding video assets may be synthesized according to the corresponding style. Specifically, a single 3D synthesis model may generate video assets in multiple styles, or a separate 3D synthesis model may be trained to generate each style.

[0161] In specific implementation, the image resources and the style in the interactive information can be input into the three-dimensional synthesis model to obtain video resources corresponding to each style; and / or, the image resources can be input into the three-dimensional synthesis model corresponding to each style to obtain the video resources output by each three-dimensional synthesis model.

[0162] Among them, different three-dimensional synthetic models correspond to different styles.

[0163] In this embodiment, a single 3D synthesis model can generate video resources in various styles. For ease of distinction, the following examples refer to this 3D synthesis model as the first 3D synthesis model. The first 3D synthesis model can include multiple branches, with different branches corresponding to different styles. In practice, image resources and the style in the interaction information can be input into the corresponding branch of the first 3D synthesis model to generate video resources that match that style. Of course, if the interaction information specifies multiple styles, the image resources can be input into multiple branches, with each branch outputting a video resource of a different style.

[0164] Alternatively, multiple styles correspond to multiple independent three-dimensional synthesis models. For the convenience of distinction, the following example refers to the three-dimensional synthesis model as the second three-dimensional synthesis model. The image resources can be input into the second three-dimensional synthesis model corresponding to the required style to obtain the video resources output by the second three-dimensional synthesis model.

[0165] Of course, in another example, the image resources can be input into a first three-dimensional synthesis model to obtain the video resources output by the corresponding branch in the first three-dimensional synthesis model, or the image resources can be input into the corresponding second three-dimensional synthesis model to obtain the video resources output by the second three-dimensional synthesis model; in this way, for the same digital resource, video resources output by multiple three-dimensional synthesis models can be obtained. In practice, the video resources output by any three-dimensional synthesis model (the first three-dimensional synthesis model and the second three-dimensional synthesis model) can be selected, such as video resources with high video quality can be selected as multimedia data.

[0166] In some examples, human-computer interaction between the server and the terminal can be achieved in various ways.

[0167] The first method involves interacting with the server through an application installed on the terminal. This allows users to achieve multimodal interaction within the application. For example, users can select the desired digital resource on the application's interface, generate interaction information, and then send it to the server through the application. After receiving the interaction information, the server can obtain the digital resource and generate multimedia data based on the digital resource.

[0168] Specifically, in this implementation, the server can receive interaction information sent by a target program located on the first terminal, wherein the target program is configured to send interaction information in response to a human-computer interaction operation performed by a user, and the interaction information carries a user account on the first terminal.

[0169] The application can be software installed on a terminal that assists a user in controlling a 3D display through human-computer interaction. In this example, the interaction information can include a user account on the first terminal, where the target program resides. The user account can refer to the account of the user logged into the first terminal.

[0170] In this way, when the interaction information instructs to play multimedia data, the interaction information can be used to instruct to play the multimedia data on the first terminal or a second terminal different from the first terminal. Specifically, when playing multimedia data, the second terminal bound to the user account can be determined first. The second terminal bound to the user account can be different from the first terminal. For example, the user account on the first terminal belongs to Mr. Li, and Mr. Li associates multiple terminals with this user account. In this way, the multiple terminals can be referred to as terminals bound to this user account.

[0171] Among them, when performing human-computer interaction operations in the target program on the first terminal, multiple second terminals bound to the user account can be controlled to play corresponding multimedia data at the same time, thereby achieving the effect of operating on one terminal and controlling multiple terminals to play multimedia data at the same time.

[0172] 6 , a schematic diagram of a display scenario is shown, including a first terminal, a server, and a second terminal, wherein a user performs human-computer interaction on a target program on the first terminal, and sends interaction information to the server through the target program on the first terminal. The server extracts multimedia data specified by the interaction information, determines a user account on the first terminal, and then determines a second terminal bound to the user account. Subsequently, the multimedia data is sent to the second terminal so that the multimedia data can be played on the second terminal.

[0173] Specifically, as shown in Figure 6, the second terminal may include multiple terminals, that is, the multimedia data can be played on multiple second terminals at the same time. The first terminal and the second terminal can be terminals of the same type, such as LED display screens, which are configured with a target program to process interactive information.

[0174] For example, the scenario shown in Figure 6 can be applied to the display of a supermarket. For example, a staff member can make gestures in front of any display screen set up in the supermarket to achieve human-computer interaction. The target program sends the gesture information to the server, and then the server can adjust the multimedia data on multiple display screens, and then achieve synchronous display of the multimedia data on multiple display screens.

[0175] It should be noted that, in this example, the interaction information is generated based on gesture operations. In some other examples, the interaction information can also be generated based on input voice, and there is no special limitation here.

[0176] In the second method, the server and the terminal can interact based on USB and HDMI. This implementation method mainly uses USB and HDMI as the main physical connection means to realize multimodal interaction of digital resources. In this way, the first terminal where the target program is located can directly establish a physical connection with the second terminal via a USB cable, and use the Bridge information communication module to identify and bind the user ID on the second terminal. After confirmation by both parties, data transmission can be carried out. Thus, the first terminal can transmit multimedia data directly to the second terminal for display via an HDMI cable. In addition, users can also interact through the gesture recognition module of the album terminal to increase the diversity of interaction.

[0177] Referring to Figure 7, another display scenario schematic diagram is shown. As shown in Figure 7, it includes a server, a first terminal, and a second terminal, wherein the first terminal and the second terminal interact via USB and HDMI, and a target program is installed on the first terminal. In this case, the user can use the target program on the first terminal for human-computer interaction, and the interaction information can be sent to the server. After the server prepares the multimedia data, it returns it to the first terminal, and then the first terminal directly sends it to the second terminal via USB and HDMI for playback.

[0178] In this scenario, the second terminal may be a display screen with only display capabilities, which may not establish a network connection with the server, but supports human-computer interaction through the first terminal and controls the display content on the second terminal through the first terminal.

[0179] Among them, in the above-mentioned first mode and second mode, the target program can be a small program, or an independent APP, or a web client program, which is not specifically limited here.

[0180] In another example, the interaction information may be obtained by an action performed by the user, which may be a gesture, body movement, head movement, etc. In this way, the condition for triggering the generation of the interaction information may be the condition that a multimedia resource on the terminal is opened. Specifically, the interaction information may be sent by the third terminal after recognizing the user's action in response to the multimedia resource being opened on the third terminal.

[0181] In this example, the third terminal may be the first terminal or the second terminal described above, without limitation. The multimedia resource on the third terminal may be a photo album on the third terminal, or a folder storing image or video data, and the operation of opening the multimedia resource may be an operation of opening the photo album. Specifically, the operation may be an operation performed by a user through an input device, such as a click operation performed by holding a mouse, or a touch operation performed on a display screen.

[0182] Among them, when it is detected that the multimedia resource is opened, the image acquisition device connected to the ring third terminal can be used to capture the user's actions through the image acquisition device. For example, continuous acquisition can be performed, such as recording a video of a certain length, etc. Therefore, the video or multiple continuous images captured by the image acquisition device can be used to identify the user's actions.

[0183] The action may be recognized by a lightweight model configured on the third terminal, or by the server, and then a recognition result may be obtained, and interaction information may be generated based on the recognition result.

[0184] It should be noted that after detecting the opening operation of the multimedia resources, the user's actions can ensure the selection of digital resources in the multimedia resources. For example, when the user extends his index finger, it indicates that image No. 1 is selected, and when he extends his index finger and middle finger, it indicates that image No. 2 is selected. Therefore, the recognition result can also include the digital resources selected by the user in the multimedia resources.

[0185] As described in the above embodiment, the digital resource selected by the user can be used as a resource for synthesizing multimedia data, or can be used to obtain a second digital resource similar to the selected digital resource.

[0186] Of course, the user can also select multimedia data to be played or shared from the multimedia resources, so that the selected multimedia data can be shared to the second terminal and played on the second terminal.

[0187] Specifically, in one example, if multimedia data is to be shared, the interaction information can include a first account and a second account, where the first account needs to share the multimedia data with the second account. In this example, the interaction information can be based on the user's voice input and the user's actions. For example, if the user inputs the voice input "share 3D video", the terminal wakes up the image acquisition device and displays the various multimedia data associated with the first account. Then, in response to the selection operation of the associated multiple multimedia data, the selected first multimedia data can be sent to the terminal where the second account is located.

[0188] Specifically, when displaying multiple multimedia data, an interface may pop up, in which the user is instructed to make gestures to select the multimedia data to be shared. Then, based on the identifier of the multimedia data selected by the user and the input voice, interaction information is generated. The server can then share the multimedia data between the first account and the second account based on the interaction information.

[0189] In this example, the various multimedia data associated with the first account can be stored on the server. When the multimedia data needs to be shared, it can be obtained from the server and returned to the terminal where the first account is located for display.

[0190] When the user selects multimedia data, the selection operation may be a gesture operation, such as moving the index finger to indicate the selection of one multimedia data, and moving the index finger and the middle finger to indicate the selection of another multimedia data.

[0191] It should be noted that the first account and the second account can be different accounts. They can be usernames. For example, if the user uses communication software, the first account and the second account can be account information set in the communication software, such as nicknames or account numbers. The first account and the second account can belong to different communication software. For example, the first account is a QQ account and the second account can be a WeChat account. In this way, multimedia data in the QQ account can be shared with the WeChat account through the server.

[0192] For example, referring to FIG8 , a schematic diagram of another display scenario is shown. As shown in FIG8 , the first terminal, the second terminal, and the server are included. The first terminal has first communication software installed, the second terminal has second communication software installed, and the first terminal and the second terminal are both large, all-in-one display devices. Specifically, the first terminal and the second terminal can be located in different locations in a supermarket. The user performs voice input on the first terminal and wishes to share multimedia data. At this time, the first terminal displays various multimedia data on the first account, collects the identifiers of the multimedia data selected by the user, generates interaction information based on the identifiers and the voice input, and feeds the interaction information back to the server.

[0193] Then, the server can obtain the second account based on the interaction information, and send the selected multimedia data to the second terminal through the second communication software, and then play the multimedia data on the second terminal.

[0194] In this way, multimedia data can be shared between different accounts based on human-computer interaction.

[0195] In another example, for some multimedia data, the style of the puzzle data can be changed through human-computer interaction. Specifically, when switching multimedia data based on the instruction of the interactive information, multiple preset styles can be displayed on the terminal. Then, in response to the selection of the multiple preset styles, the second multimedia data can be converted into third media data that matches the target style; the target style is the selected preset style.

[0196] The identifier of the determined second multimedia data may be included in the interaction information. In this example, the interaction information may be generated based on user input, such as voice or text. The interaction information may include the identifier of the second multimedia data to be converted to a certain style, as well as a keyword indicating the need for style switching. Thus, when the server determines that a style switch is necessary, it may display a variety of preset styles for the user to select.

[0197] The multiple preset styles displayed may be displayed using text information, such as a warm color style or a black, white, and gray style, or may be displayed as a preview view of the preset styles. When the multiple preset styles are displayed, a user selection operation of the preset styles may be detected. The selection operation may be a gesture operation, a voice operation, or a click operation on a preset style. The selected preset style is referred to as a target style, and the server may convert the second multimedia data into third multimedia data that conforms to the target style.

[0198] When converting the second multimedia data into third multimedia data that conforms to the target style, if the second multimedia data is multimedia data synthesized based on digital resources, the digital resources synthesized into the second multimedia data can be re-input into the corresponding three-dimensional synthesis model to generate third multimedia data that conforms to the target style. If the second multimedia data is 2D type multimedia data and needs to be converted into 3D type multimedia data, the second multimedia data can be input into the 2D-to-3D model to obtain three-dimensional third multimedia data.

[0199] After obtaining the third multimedia data, the third multimedia data may be played, so that the user can convert the style of the displayed multimedia data based on human-computer interaction.

[0200] In another example, switching multimedia data can refer to switching the terminal displaying the multimedia data. For example, switching multimedia data playback from one terminal to another. In this example, the interaction information can include a fourth terminal to be displayed. Therefore, the fourth terminal needs to play the multimedia data, meaning that the multimedia data needs to be switched from the current terminal to the fourth terminal.

[0201] Specifically, when switching to a fourth terminal for playback, the multimedia data can be converted into multimedia data compatible with the fourth terminal. When switching multimedia data, the fourth multimedia data currently being played by the terminal that triggered the interaction information can be obtained, and the display type of the display screen configured for the fourth terminal can be obtained, and the current type includes a three-dimensional display type and a two-dimensional display type; and based on the display type and the type of the fourth multimedia data, the fourth multimedia data can be displayed on the fourth terminal.

[0202] In this example, the fourth terminal may be configured with a display screen, or the fourth terminal itself may be a display screen. The type of the fourth terminal may include a three-dimensional display type and a two-dimensional display type. Specifically, when the fourth terminal is configured with naked-eye 3D display hardware, such as the light field display device described above, the type of the fourth terminal may be determined to be a three-dimensional type. When the display screen of the fourth terminal is not configured with 3D display hardware, its type is a two-dimensional type.

[0203] The fourth multimedia data may be processed based on the difference between the display type of the display screen configured for the fourth terminal and the type of the fourth multimedia data to ensure that the type of the fourth multimedia data and the display type are consistent. For example, if the display type of the fourth terminal is 3D and the type of the fourth multimedia data is 3D, the fourth multimedia data does not need to be processed and can be directly displayed on the fourth terminal. If the display type of the fourth terminal is 2D and the type of the fourth multimedia data is 3D, the fourth multimedia data may be converted to 2D data before display; if the display type of the fourth terminal is 3D and the type of the fourth multimedia data is 2D, the fourth multimedia data may be converted to 3D data before display.

[0204] By adopting the technical solution of this example, when switching multimedia data, the server can automatically convert the type of the fourth multimedia data according to the terminal to be displayed, that is, the display type of the terminal after switching, so that the fourth multimedia data can adapt to the display type of the terminal after switching, and then in the human-computer interaction display, the displayed multimedia data can adapt to the display terminal.

[0205] Several examples are given below to illustrate the above embodiment:

[0206] Example 1,

[0207] 9 , a flow diagram of the display method of Example 1 is shown. As shown in FIG9 , the method includes a server (cloud), a first terminal, and a second terminal. The first terminal is a mobile phone, and a target program is running on the first terminal. The second terminal is a large LED display screen and is equipped with a naked-eye 3D display component. The first terminal and the second terminal can communicate wirelessly, such as via WIFI or Bluetooth, and the target program is a small program on the mobile phone.

[0208] In this example, the user can perform human-computer interaction on the mobile device. In practice, after the user opens the target program, he can enter text information in the target program, such as entering "play the virtual three-dimensional animation of the pig". The resource request module of the target program can send the text information as interactive information to the server. At this time, the server can process the text information and input the processed text information into the resource generation model. The resource generation model outputs multiple two-dimensional images of the pig. Then, the multiple two-dimensional images are sent to the three-dimensional synthesis model. The three-dimensional synthesis model converts the two-dimensional images into three-dimensional images, and sends the multiple frames of three-dimensional images as three-dimensional videos to the second terminal and plays them on the second terminal.

[0209] Among them, when the three-dimensional synthesis model generates a three-dimensional video, it can generate three-dimensional videos of multiple different styles and send the three-dimensional videos of multiple different styles to the second terminal. Then, the user can select the three-dimensional video of the required style to play. During this process, the second terminal can display preview images of the multiple three-dimensional videos. Then, the second terminal can use the configured image acquisition device to capture images of the user's gestures and send the captured images to the first terminal. Then, the first terminal feeds back the recognition results to the second terminal, so that the second terminal plays the three-dimensional video selected by the user.

[0210] In the implementation scheme of this example 1, the user can complete the generation, switching and control of the three-dimensional video displayed on the LED display screen through human-computer interaction on the mobile phone, and can also perform simple human-computer interaction on the second terminal, the LED display screen, to achieve selective playback of the three-dimensional content to be displayed on the display screen.

[0211] Example 2

[0212] 10 , a flow chart of the display method of Example 2 is shown. As shown in FIG10 , the method includes a server, a first terminal, and a web terminal. Both the web terminal and the first terminal are communicatively linked to the server (cloud), and the first terminal can be an integrated display machine. A user can select a digital resource through a web terminal (e.g., a browser), and the resource will be uploaded to a data storage module of the server. Then, the server will find a digital resource similar to the digital resource (referred to as the second digital resource above) based on the uploaded digital resource. These digital resources include audio resources and image resources. A three-dimensional video resource is then obtained through a 2D to 3D conversion model (a three-dimensional synthesis model), and audio adapted to the video resource is obtained after processing by an audio and video synthesis display module. The audio and video resources are synthesized into a three-dimensional video. Then, the server sends the three-dimensional video to the first terminal for display.

[0213] Among them, on the first terminal side, the first terminal can store the three-dimensional video in a multimedia resource, such as a photo album. In practice, the first terminal can play the multimedia data selected by the user in the multimedia resource through the user gesture captured by the image capture device.

[0214] The implementation scheme of this Example 2 can control the display of 3D content on the terminal by human-computer interaction through the browser. The terminal can also store the 3D video and selectively play the stored multimedia data by recognizing the user's gestures.

[0215] Example 3

[0216] The system includes a server and a third terminal, wherein the third terminal is connected to the server and is an all-in-one display device. The third terminal can be configured with an image acquisition module and a lightweight gesture recognition module. When the multimedia resources on the third terminal are opened, the image acquisition module is awakened and the user's gesture is collected. Then, the gesture recognition module recognizes the gesture, generates interaction information based on the recognition result, and sends the interaction information to the server.

[0217] In this example 3, after the multimedia resource is opened, the user can select one of the multimedia data through a gesture. Then, the server can convert the style and type of the multimedia data based on the selected multimedia data and send it to the third terminal for playback on the third terminal.

[0218] To sum up, when adopting the display method of this example, at least one of the operations of generating, playing, switching, storing, querying and sharing multimedia data can be performed directly based on the user's interactive operation. For example, at least three-dimensional multimedia data can be generated based on human-computer interactive operations. Therefore, in 3D display technology, the indicated three-dimensional video and / or three-dimensional image can be automatically synthesized according to the acquired interactive operations, so that human-computer interaction can be integrated into the display control and content generation of the three-dimensional display, thereby providing users with a new, intuitive and efficient interactive experience.

[0219] Based on the same inventive concept, a display system is also proposed. Referring to Figure 11, a schematic diagram of the layout structure of a display system is shown. As shown in Figure 11, the display system may include a server and a terminal connected to the server, wherein the terminal is configured with a three-dimensional display screen.

[0220] Specifically, the server can be connected to multiple terminals, each of which can be configured with a three-dimensional display screen. The models and types of the multiple terminals may not be exactly the same. For example, they may include all-in-one display machines, large LED display screens, large naked-eye 3D display screens, mobile terminals, computers, etc.

[0221] In this embodiment, the server can be used to obtain interaction information and perform target operations on multimedia data based on the interaction information; wherein the target operation includes at least one of generating, playing, switching, storing, querying, and sharing multimedia data; the interaction information is generated based on at least one of a user input voice, a performed action, and an input text; wherein the digital resource includes at least one of an image resource, an audio resource, and a video resource;

[0222] The terminal can be used to play multimedia data after executing the target operation.

[0223] Referring to FIG12 , an exemplary display system structure diagram of the present embodiment is provided. As shown in FIG12 , the software and hardware settings required for the display system, as well as the layout structure between the software and hardware, are listed in detail. Among them, the system is mainly composed of four parts: mac / windows software, server, terminal, and related hardware, which realizes efficient collaboration between the terminal local and the server. The software side is responsible for processing resource requests, device binding, and device control, and uses a speech-to-text (STT) model and a text generation model similar to GPT for high-quality speech recognition and text processing and generation. This provides users with an intuitive and convenient way of interaction, and through generation technology, makes human-computer interaction smoother.

[0224] Specifically, the server can be configured with a 3D synthesis model and a resource generation model. The resource generation model can be a diffusion image generation model, which generates 3D multimedia data based on digital resources. The resource generation module generates virtual digital resources, such as audio and image-based virtual digital resources, based on input text information. The server also includes a communication module, a data storage module, a resource module, a monitoring and alarm module, and an audio and video synthesis display module, providing comprehensive support for the system and ensuring its stable operation.

[0225] Among them, the communication module is used to ensure communication with the terminal, the data storage module can be used to store digital resources and multimedia data, the resource module can be used to obtain digital resources based on interactive information, and connect with multiple data sources to pull the required digital resources from the data source, the monitoring and alarm module can be used to detect the operating status of the system and issue an alarm in the event of an operating failure or overload; the audio and video synthesis and display module is used to synthesize video resources and audio resources to obtain video-type multimedia data.

[0226] The terminal may be configured with a lightweight gesture recognition module for recognizing gestures performed by the user, thereby selecting and playing multimedia data on the terminal based on the recognition result, and generating interaction information based on the recognition result and sending it to the server.

[0227] In terms of hardware, the display system includes a 7.9-inch OLED screen that can be connected to a terminal or used directly as an all-in-one display to provide a high-quality visual experience. Key hardware components such as an infrared camera, mini speakers, grating lenses, and a SoC chip ensure the system's multimodal human-computer interaction capabilities.

[0228] All software, servers, and terminals can communicate securely via HTTPS. Furthermore, the software can be connected directly to the terminal via a USB cable. For example, if the software is a small program, the program can run on a mobile terminal, which can then communicate with the terminal equipped with a display via a USB cable. The bridge information communication module can identify and bind device user IDs, ensuring system security.

[0229] In this embodiment, the system is also configured with a target program, which is set on the terminal or on an electronic device connected to the terminal; wherein, when the target program is running, it is used to implement the following steps: based on the target operation input by the user, sending the interaction information to the server; wherein, the target operation includes at least one of voice operation, gesture operation, and text input operation.

[0230] In this embodiment, the target program can be a small program or a web client program, which can be run on a mobile terminal. In this case, the target program can respond to the target operation input by the user by sending interaction information to the server. The server can then perform corresponding operations on the multimedia data on the terminal based on the interaction information.

[0231] In some examples, the server is configured with a resource generation model, which is configured to generate virtual digital resources based on input text information; wherein the virtual data resources include audio type resources and / or image type resources; wherein the resource generation model is trained using multiple text samples and virtual digital resource samples corresponding to each text sample as training samples.

[0232] The process of generating virtual digital resources by the resource generation model in this example can refer to the description of the embodiment of the above-mentioned display method, and will not be described in detail here.

[0233] In some embodiments, the server is further configured with at least one three-dimensional synthesis model, wherein the three-dimensional synthesis model is used to generate three-dimensional image resources based on input image resources; wherein different three-dimensional synthesis models correspond to different styles.

[0234] The process of generating multimedia data from the three-dimensional synthetic model in this example can refer to the description of the embodiment of the above-mentioned display method, and will not be described in detail here.

[0235] In some examples, the terminal is configured with a motion recognition module, and the motion recognition module is configured to, in response to an operation of opening a multimedia resource on the third terminal, recognize a user's motion and then send the interaction information;

[0236] The action includes at least one of a gesture action, a body action, and a head action.

[0237] The process of the action recognition module performing gesture recognition and uploading interaction information may refer to the description of the embodiment of the above-mentioned display method, and will not be elaborated here.

[0238] In some examples, the server is further configured with a database storing a variety of digital resources; the terminal is further configured with a resource library including a variety of multimedia resources; wherein the digital resources are derived from the resource library and / or from the database.

[0239] Based on the same inventive concept, the present disclosure also provides a display device that can be configured on a server and a terminal to jointly implement the display method described in the above example. Specifically, referring to the figure, a schematic diagram of the structure of a display device is shown, which may include:

[0240] An information acquisition module is used to acquire interaction information, where the interaction information is generated based on a human-computer interaction operation performed by a user;

[0241] a response module, configured to perform a target operation based on the interaction information, wherein the target operation includes at least one of generating, playing, switching, storing, querying, and sharing multimedia data;

[0242] Generating multimedia data includes: obtaining digital resources to be synthesized based on the interactive information, and generating three-dimensional multimedia data based on the digital resources; the digital resources include at least one of audio resources, image resources and video resources.

[0243] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0244] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, commodity, or device that includes the element.

[0245] The above is a detailed introduction to a display method and system provided by the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.

[0246] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0247] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

[0248] References herein to "one embodiment," "an embodiment," or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Furthermore, please note that instances of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.

[0249] In the description provided herein, numerous specific details are described. However, it is understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.

[0250] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present disclosure may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0251] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.

Claims

1. A display method, characterized in that: The method comprises: Acquiring interaction information, where the interaction information is generated based on a human-computer interaction operation performed by a user; performing a target operation based on the interaction information, the target operation comprising at least one of generating, playing, switching, storing, querying, and sharing multimedia data; Generating multimedia data includes: obtaining digital resources to be synthesized based on the interactive information, and generating three-dimensional multimedia data based on the digital resources; the digital resources include at least one of audio resources, image resources and video resources.

2. The display method according to claim 1, wherein: The human-computer interaction operation includes at least one of a voice input operation, a body gesture operation, and a text input operation.

3. The display method according to claim 1, wherein: The step of obtaining the digital resource to be synthesized based on the interaction information includes: Parsing the interaction information to obtain first text information, where the first text information is used to indicate a location and / or identifier of a digital resource; The digital resource is acquired based on the first text information.

4. The display method according to claim 3, wherein: The interaction information is generated according to an opening operation of a multimedia resource on a terminal, and the acquiring of the digital resource based on the first text information includes: Determining a first data source where the data resource is located based on the first text information; Acquire a first digital resource from the first data source, and acquire a second digital resource from a second data source different from the first data source; wherein a similarity between the second digital resource and the first digital resource is greater than a preset similarity; The first digital resource and / or the second digital resource are used as the digital resources to be synthesized.

5. The display method according to claim 1, wherein: The step of obtaining the digital resource to be synthesized based on the interaction information includes: Parsing the interaction information to obtain text information, where the text information is used to at least indicate content of the multimedia data; Inputting the text information into at least one resource generation model to obtain a virtual digital resource generated by the resource generation model; wherein the resource generation model is trained using a plurality of text samples and virtual digital resource samples corresponding to each text sample as training samples; The virtual data resources include audio type resources and / or image type resources.

6. The display method according to claim 1, wherein: The generating the multimedia data based on the digital resource includes: Acquire an image resource belonging to an image type from the digital resources, and acquire an audio resource belonging to an audio type from the digital resources; Generate video resources based on the image resources; The video resource and the audio resource are synthesized to obtain the multimedia data.

7. The display method according to claim 6, wherein: The generating of the video resource based on the image resource includes: Inputting the image resource into a three-dimensional synthesis model to obtain at least one video resource output by the three-dimensional synthesis model; The three-dimensional synthesis model is obtained by training with a variety of two-dimensional image samples as training samples and three-dimensional multimedia data corresponding to the two-dimensional image samples as supervision.

8. The display method according to claim 7, wherein: The interactive information includes at least one style of multimedia data; and inputting the image resource into the three-dimensional synthesis model includes: Inputting the image resource and the style in the interaction information into the three-dimensional synthesis model to obtain video resources corresponding to each style; And / or, the image resource is input into a three-dimensional synthesis model corresponding to each of the styles to obtain a video resource output by each of the three-dimensional synthesis models, wherein different three-dimensional synthesis models correspond to different styles.

9. The display method according to claim 1, wherein: The acquiring of interaction information includes: receiving interaction information sent by a target program located on a first terminal, wherein the target program is configured to send the interaction information in response to a human-computer interaction operation performed by a user, and the interaction information carries a user account on the first terminal; The playing of the multimedia data includes: Based on the user account, determining a second terminal bound to the user account; The multimedia data is sent to the second terminal so as to play the multimedia data on the second terminal; wherein the second terminal is the first terminal or a terminal different from the first terminal.

10. The display method according to claim 1, wherein: The interaction information is sent by the third terminal after identifying the user's action in response to the operation of opening the multimedia resource on the third terminal; wherein the user's action includes at least one of a gesture action, a body action, and a head action.

11. The display method according to claim 1, wherein: The interaction information carries the first account and the second account, and the sharing of the multimedia data includes: displaying a variety of multimedia data associated with the first account; In response to a selection operation on the associated multiple multimedia data, the selected first multimedia data is sent to the terminal where the second account is located.

12. The display method according to claim 1, wherein: The interaction information carries an identifier of the second multimedia data to be converted, and the switching of the multimedia data includes: Display multiple preset styles; In response to a selection operation of the plurality of preset styles, the second multimedia data is converted into third multimedia data conforming to a target style; wherein the target style is the selected preset style.

13. The display method according to claim 1, wherein: The interaction information carries a fourth terminal to be displayed, and the switching of the multimedia data includes: Acquiring fourth multimedia data currently played by the terminal that triggers the interaction information; Obtaining a display type of a display screen configured for the fourth terminal, where the display type includes a three-dimensional display type and a two-dimensional display type; Based on the display type and the type of the fourth multimedia data, the fourth multimedia data is displayed on the fourth terminal.

14. A display system, characterized in that: The system includes a server and a terminal connected to the server, wherein the terminal is equipped with a three-dimensional display screen, wherein: The server is configured to obtain interaction information and perform a target operation on the multimedia data based on the interaction information; wherein the target operation includes at least one of generating, playing, switching, storing, querying, and sharing the multimedia data; the interaction information is generated based on at least one of a voice input, an action performed, and a text input by the user; and the digital resource includes at least one of an image resource, an audio resource, and a video resource; The terminal is used to play the multimedia data after performing the target operation.

15. The display system according to claim 14, wherein: The system is further configured with a target program, which is set on the terminal or on an electronic device connected to the terminal; wherein the target program, when running, is used to implement the following steps: Sending the interaction information to the server based on the target operation input by the user; The target operation includes at least one of a voice operation, a gesture operation, and a text input operation.

16. The display system according to claim 14, wherein: The server is configured with a resource generation model, and the resource generation model is configured to generate virtual digital resources based on input text information; wherein the virtual data resources include audio type resources and / or image type resources; The resource generation model is obtained by training using a variety of text samples and virtual digital resource samples corresponding to each text sample as training samples.

17. The display system according to claim 14, wherein: The server is configured with at least one three-dimensional synthesis model, wherein the three-dimensional synthesis model is used to generate a three-dimensional image resource based on an input image resource; Among them, different three-dimensional synthetic models correspond to different styles.

18. The display system according to claim 14, wherein: The terminal is provided with a motion recognition module, wherein the motion recognition module is configured to, in response to an operation of opening a multimedia resource on the third terminal, recognize a user's motion and then send the interaction information; The action includes at least one of a gesture action, a body action, and a head action.

Citation Information

Patent Citations

  • Real 3D virtual simulation interaction method and system

    CN108831216A

  • Method and device for human-computer interaction

    CN112181127A

  • Interaction method, device and equipment of self-service terminal and storage medium

    CN113900565A

  • Data interaction system

    CN117610093A

  • Display method and display system

    CN118312071A