Method and apparatus for providing media content, and electronic device and storage medium

By selecting dialogue items with virtual objects in a virtual scene and generating media resources associated with the virtual scene, the problem of high repetition rate in traditional media content recommendations is solved, the provision of personalized media content is realized, and the user interaction experience is improved.

WO2025112902A1PCT designated stage expired Publication Date: 2025-06-05BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/123069
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-30
Filing Date
2024-09-30
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The repetition rate between media content recommended by traditional electronic devices for different users is high, which cannot meet the personalized needs of different users for media content and affects the user's experience.

Method used

Personalized media content is provided by receiving the selection of dialogue items in the conversation interaction between the target user and the virtual object in the virtual scene and generating media resources associated with the virtual scene based on the selected dialogue items.

Benefits of technology

It improves users' interactive experience in virtual scenes, provides personalized media content, and meets the personalized needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024123069_05062025_PF_FP_ABST
    Figure CN2024123069_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present embodiment relate to a method and apparatus for providing media content, and an electronic device and a storage medium. The method comprises: receiving a selection of a group of dialogue items provided for at least one round of dialogue interaction in at least one round of dialogue interaction between a target user and a virtual object in a virtual scene; and providing target media content, wherein the target media content is generated on the basis of media resources corresponding to at least one dialogue item among the group of dialogue items, the media resources being associated with the virtual scene. On the basis of the embodiments of the present disclosure, a target user can acquire, by means of selecting corresponding dialogue items, target media content generated on the basis of media resources corresponding to the dialogue items selected by the target user in dialogue interaction with a virtual object in a virtual scene.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, electronic device and storage medium for providing media content

[0001] This application claims priority to the Chinese invention patent application entitled “Method, device, electronic device and storage medium for providing media content” filed on November 30, 2023, with application number 202311632633.X. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to a method, apparatus, electronic device, and storage medium for providing media content. Background Art

[0003] With the development of computer technology, various electronic devices can greatly enrich people's daily life. For example, people can use electronic devices to carry out various interactions.

[0004] In some interactive scenarios, people can, for example, control various virtual characters in the virtual scene to complete the interaction. How to improve people's interactive experience in the virtual scene is a key issue.

[0005] Summary of the Invention

[0006] In a first aspect of the present disclosure, a method for providing media content is provided. The method comprises: receiving a selection of a set of dialogue items provided in at least one round of dialogue interaction between a target user and a virtual object in a virtual scene; and providing target media content, where the target media content is generated based on a media resource corresponding to at least one dialogue item in the set of dialogue items, the media resource being associated with the virtual scene.

[0007] In a second aspect of the present disclosure, a device for providing media content is provided. The device includes: a receiving module configured to receive, during at least one round of conversational interaction between a target user and a virtual object in a virtual scene, a selection of a set of dialogue items provided in the at least one round of conversational interaction; and a providing module configured to provide target media content, where the target media content is generated based on a media resource corresponding to at least one dialogue item in the set of dialogue items, the media resource being associated with the virtual scene.

[0008] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0009] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0010] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0012] FIG1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0013] FIG2 illustrates a flow chart of a process for providing media content according to some embodiments of the present disclosure;

[0014] 3A to 3E illustrate example interaction interfaces according to some embodiments of the present disclosure;

[0015] FIG4 shows a schematic structural block diagram of an apparatus for providing media content according to certain embodiments of the present disclosure;

[0016] FIG5 shows a block diagram of an electronic device capable of implementing various embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0020] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the relevant laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out under the premise that the user is aware of and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. The specific notification and / or authorization method may vary according to actual conditions and application scenarios, and the scope of the present disclosure is not limited in this respect.

[0021] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.

[0022] Traditionally, electronic devices recommend a high repetition rate of media content to different users, which may not meet the personalized needs of different users and also affect the user experience.

[0023] Embodiments of the present disclosure provide a media content provision solution. According to various embodiments of the present disclosure, during at least one round of conversational interaction between a target user and a virtual object in a virtual scene, a selection of a set of dialogue items provided in the at least one round of conversational interaction can be received; and target media content can be provided, where the target media content is generated based on a media resource corresponding to at least one dialogue item in the set of dialogue items, the media resource being associated with the virtual scene.

[0024] According to the embodiments of the present disclosure, a target user can obtain personalized target media content by interacting with virtual objects in a virtual scene. In addition, the target media content is generated based on the media resources associated with the virtual scene, which helps the user perceive the virtual scene more effectively, thereby improving the user's interactive experience.

[0025] Sample Environment

[0026] FIG1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG1 , the example environment 100 may include a terminal device 110 .

[0027] In this example environment 100, a terminal device 110 may run an application 120 that supports a virtual scene. Application 120 may be any suitable type of application for presenting a virtual scene, and examples thereof may include, but are not limited to, simulation applications, emulation applications, gaming applications, virtual reality applications, augmented reality applications, and the like, although the embodiments of the present disclosure are not limited in this respect. User 140 may interact with application 120 via terminal device 110 and / or its attached devices.

[0028] In the environment 100 of Figure 1 , if the application 120 is active, the terminal device 110 can present an interface 150 associated with the virtual scene through the application 120. The interface 150 can present at least one screen associated with the virtual scene. The at least one screen can include a screen associated with a virtual object corresponding to the current user, a screen associated with virtual objects corresponding to other users, a screen corresponding to a non-player character, a screen associated with a location in the virtual scene, and the like. For example, the interface 150 can be a game application interface to present the corresponding game scene. Alternatively, the interface 150 can be another appropriate type of interactive interface that can support users controlling virtual objects in the interface to perform corresponding actions in the virtual scene.

[0029] In some embodiments, the terminal device 110 communicates with the server 130 to enable the provision of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a handheld computer, a portable game terminal, a VR / AR device, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as a "wearable" circuit, etc.).

[0030] Server 130 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server 130 can include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, and the like. Server 130 provides backend services for application 120 that supports virtual scenarios in terminal device 110.

[0031] A communication connection may be established between the server 130 and the terminal device 110. The communication connection may be established via a wired or wireless method. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus (USB) connection, a Wireless Fidelity (WiFi) connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the server 130 and the terminal device 110 may implement signaling interaction via the communication connection between the two.

[0032] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0033] Example Process

[0034] FIG2 shows a flow chart of a process 200 for providing media content according to some embodiments of the present disclosure. The process 200 may be implemented at the terminal device 110. The process 200 is described below with reference to FIG1.

[0035] In block 210 , the terminal device 110 receives a selection of a set of dialogue items provided in at least one round of dialogue interaction between a target user and a virtual object in a virtual scene.

[0036] In some embodiments, the virtual scene may be a game scene, a simulation scene, a virtual reality scene, etc. The virtual object may be an appropriate interactive object set in the virtual scene, wherein the interactive object is a non-player character, such as a virtual character in the virtual scene, etc.

[0037] The specific process of block 210 will be described below in conjunction with Figures 3A to 3E. It should be understood that although Figures 3A to 3E are described using a game scene as an example, this is merely exemplary.

[0038] As shown in Figure 3A, the terminal device 110 can present an interface 300A, which can be used to present relevant information of a virtual scene (e.g., a game scene). Exemplarily, the interface 300A can include a virtual object 310, such as a non-player character in the virtual scene.

[0039] Furthermore, the terminal device 110 may also provide a dialogue control 320 to support dialogue interaction between the target user and the virtual object 310. For example, such a dialogue control 320 may be provided based on a dialogue request between the target user and the virtual object. Alternatively, such a dialogue control 320 may be automatically provided in the virtual scene based on the satisfaction of specific trigger conditions. Such trigger conditions may include, but are not limited to, any appropriate plot conditions, time conditions associated with the virtual scene, character conditions corresponding to the player character in the virtual scene, and the like.

[0040] 3A , in the dialogue control 320 , the terminal device 110 may present a first sentence 330 provided by the virtual object 310. For example, the first sentence 330 may be “Hello, would you like me to draw a picture for you? Which city do you like best on this continent?”

[0041] It should be understood that, although the first sentence 330 is provided in text form in FIG. 3A , the first sentence 330 may also be provided to the target user in other forms such as audio and video.

[0042] Furthermore, after the first statement 330 is provided, if the preset dialogue item provision conditions are met, the terminal device 110 may provide a group of candidate dialogue items 340-1 to 340-3 (individually or collectively referred to as dialogue items 340) as shown in FIG3B . The preset dialogue item provision conditions in the embodiment of the present disclosure may be that the time length from the time the first statement 330 is provided to the current time reaches a preset time length threshold, or a preset operation of the target user is received (such as the target user clicks on the dialogue item selection control on the display page of the terminal device 110), etc. Such a group of candidate dialogue items 340 may indicate candidate responses to the first statement 330 provided by the virtual object 310. It should be noted that different dialogue items in such a group of candidate dialogue items 340 indicate different candidate responses.

[0043] Furthermore, the terminal device 110 may receive the target user's selection of one or more dialogue items in the group of dialogue items 340 , thereby completing a round of dialogue interaction with the virtual object 310 .

[0044] In addition, it should be understood that in some scenarios, the target user may also initiate a conversation. That is, the target user may, for example, start a conversation interaction with the virtual object 310 by selecting a conversation item from a plurality of conversation items.

[0045] Thus, a round of conversational interaction can always include a conversation item selected by the user and a statement provided by a virtual object. This can include the user selecting a conversation item and the virtual object providing a response statement based on the conversation item; or the virtual object providing a question statement and the user selecting a conversation item as a response statement.

[0046] Taking FIG. 3A and FIG. 3B as examples, the dialogue interaction process shown in FIG. 3A and FIG. 3B can be regarded as one round of dialogue interaction.

[0047] In some embodiments, the target user may also engage in multiple rounds of conversational interaction with virtual object 310. For example, as shown in FIG3C , in interface 300C, virtual object 310 may further provide a second statement 350. Similarly, after second statement 350 is provided, if preset conditions for providing conversation items are met, terminal device 110 may provide a set of candidate conversation items 360-1 through 360-3 (individually or collectively referred to as conversation items 360), as shown in FIG3D . The preset conditions for providing conversation items in the disclosed embodiments may include the time length from the time second statement 350 was provided to the current time reaching a preset time threshold, or receiving a preset action from the target user (e.g., the target user clicking a conversation item selection control on a display screen of terminal device 110). This set of candidate conversation items 360 may indicate candidate responses to second statement 350 provided by virtual object 310. It should be noted that different conversation items in this set of candidate conversation items 360 indicate different candidate responses.

[0048] Furthermore, the terminal device 110 may receive the target user's selection of one or more dialogue items in the group of dialogue items 360 , thereby completing a second round of dialogue interaction with the virtual object 310 .

[0049] It should be understood that such a dialog interaction process may be performed for a predetermined number of rounds, which may be a preset number or may be determined based on the dialog item selected by the target user.

[0050] The above describes the basic interaction process of dialogue interaction. The following describes the process of generating sentences and dialogue items provided by virtual objects. It should be understood that the generation process can be performed by an appropriate generation device (e.g., terminal device 110, server 130, and / or other appropriate electronic devices or combinations thereof in FIG1 ).

[0051] In some embodiments, the first statement 330 for initiating a conversational interaction and / or the set of candidate conversation items for initiating a conversational interaction may include preset content. For example, using a non-player character as an example, different users may obtain the same first statement 330 and / or conversation items when they begin interacting with the same non-player character in a virtual scene.

[0052] Alternatively, the generation device may also generate the first sentence 330 and / or the corresponding dialogue item based on at least one descriptive information item associated with the target user in the virtual scene (e.g., when the target user selects the dialogue item to trigger the dialogue interaction). For example, the generation device may generate the corresponding first sentence 330 and / or the dialogue item based on information (e.g., name, personality, occupation, etc.) of the character of the player controlled by the target user.

[0053] Alternatively or additionally, the generating device may also generate the first sentence 330 and / or the dialogue item based on the historical interactions between the target user and the virtual object. Such historical interactions may include historical dialogue interactions, or other types of interactions in the virtual scene (e.g., combat interactions, team interactions, transaction interactions, etc.).

[0054] In this way, the embodiments of the present disclosure can provide more personalized dialogue interactions for different users, thereby improving the user interaction experience.

[0055] Furthermore, for multi-round conversational interactions, the results of the previous round of conversational interaction can also be used to guide the generation of sentences and / or conversation items for the next round of conversation. For example, using FIG3B as an example, if the target user selects conversation item 340-1, the generation device can generate a second sentence 350 and / or a set of candidate conversation items 360 for the second sentence, as shown in FIG3C , based on the selected conversation item 340-1.

[0056] In some embodiments, each set of dialogue items provided in a multi-round dialogue interaction may include multiple dialogue items corresponding to different preset media resources. Taking Figures 3B and 3D as examples, dialogue items 340-1, 340-2, and 340-3 may correspond to different media resources, and dialogue items 360-1, 360-2, and 360-3 may correspond to different media resources.

[0057] For example, such media resources may include different first media resources, where the first media resources are media resources associated with virtual locations included in the virtual scene (e.g., virtual location 1, virtual location 2, and virtual location 3, etc.), or different second media resources, where the second media resources are media resources associated with virtual props included in the virtual scene (e.g., prop 1, prop 2, and prop 3, etc.). Such media resources may be provided to a generation device as guidance for generating dialogue items 340 and / or 360.

[0058] In some embodiments, the generation device may, for example, utilize an appropriate generative model to generate sentences and dialogue items in a conversational interaction. Such a generative model may, for example, include any appropriate machine learning model to process input information according to the process discussed above to generate corresponding sentences and / or dialogue items. This disclosure is not intended to limit the specific structure and type of the generative model.

[0059] It should be understood that the conversational interaction discussed above may last for one round or multiple rounds.

[0060] In addition, before ending the dialogue interaction with the target user, the terminal device 110 can also output a prompt message to prompt the target user to wait for the output of media content. For example, as shown in Figure 3E, the terminal device 110 can output a prompt message 370, such as outputting a prompt message "I already know the picture you imagine. Please wait for a while to pick up the painting."

[0061] In block 220 , the terminal device 110 provides target media content, where the target media content is generated based on a media resource corresponding to at least one dialog item in a group of dialog items, and the media resource is associated with the virtual scene.

[0062] In the disclosed embodiment, after the target user selects the corresponding conversation item, the terminal device 110 or the server 130 can obtain the corresponding media resources based on the conversation item selected by the target user. Each conversation item corresponds to a different preset media resource. The media resources may include media resources in the form of images, audio, video, etc.

[0063] Media resources may include a first media resource associated with a virtual location within a virtual scene. The virtual location may be a real place or city within the virtual scene. Media resources may also include a second media resource associated with a virtual character within the virtual scene. Media resources may also include a third media resource associated with a virtual item within the virtual scene. For example, if the virtual scene is a game scene, the virtual character may be a hero selected or owned by the target user, and the virtual item may be a game prop (such as a weapon) selected or owned by the target user.

[0064] The following describes in detail the process of obtaining corresponding media resources based on a conversation item. For ease of description, the following takes the example of the terminal device 110 obtaining corresponding media resources based on a conversation item.

[0065] In the disclosed embodiments, a conversation item may correspond to a single media resource or a set of candidate media resources. When a conversation item corresponds to a set of candidate media resources, the set includes multiple images associated with the virtual scene, each corresponding to a different perspective. For example, the set of candidate media resources corresponding to conversation item A includes a distant image corresponding to location A, a mid-range image corresponding to location A, a close-up image corresponding to location A, and so on.

[0066] If a conversation item corresponds to a candidate media resource, then after obtaining the target user's selection of the conversation item, the terminal device 110 may directly use the unique media resource corresponding to the conversation item as the media resource selected by the target user.

[0067] If a conversation item corresponds to a group of candidate media resources, in some embodiments, the terminal device 110 may randomly select a candidate media resource from the group as the target media resource after obtaining the target user's selection of the conversation item. To ensure that the generated target media content better meets the needs of the target user, in other embodiments, after obtaining the target user's selection of the conversation item, the terminal device 110 may also display another conversation item group on the display page, wherein the other conversation item group includes multiple other conversation items.

[0068] In some embodiments, the number of other conversation items included in the other conversation item group is equal to the number of candidate media resources corresponding to the conversation item selected by the target user, and each other conversation item corresponds to one candidate media resource. The target user can select at least one other conversation item on the displayed page, for example, by clicking a control corresponding to the at least one other conversation item. Terminal device 110 can receive the target user's selection of the at least one other conversation item and determine the candidate media resource corresponding to the at least one other conversation item as the target media resource.

[0069] Furthermore, the terminal device 110 or the server 130 may generate target media content based on the media resources selected by the target user. The following describes a process of generating target media content based on the target media resources.

[0070] In some embodiments, the terminal device 110 and / or the server 130 can generate the target media content to be provided by processing one or more media resources determined based on the selection of the target user. As an example, the terminal device 110 and / or the server 130 can generate the target media content to be provided by adjusting, modifying, or combining the target media resources in any appropriate manner.

[0071] In some other embodiments, the target media content may also be generated using a first model, for example. It should be understood that such a first model may include, for example, any appropriate machine learning model that can be trained to perform the task of generating media content, and this disclosure is not intended to be limited thereto.

[0072] In some embodiments, the terminal device 110 and / or the server 130 may generate input information for the first model based on the determined multiple media resources. Taking the first model as an image generation model as an example, the terminal device 110 and / or the server 130 may provide the image generation model with the determined one or more media resources so that the first model can generate corresponding media content based on the one or more media resources.

[0073] In some other embodiments, the input information may further include prompt words for guiding the generation of the first model. In some embodiments, such prompt words may include preset prompt words.

[0074] In yet other embodiments, such prompt words may also be constructed based on one or more dialogue items selected by the target user. For example, in some scenarios, during the target user's dialogue interaction with a virtual object in a virtual scene, the terminal device 110 may also provide the target user with a set of scenario selection dialogue items on a display page, each of which indicates a type of scenario information, such as combat scenario information, scenery viewing scenario information, transaction scenario information with a merchant, and so on.

[0075] The target user can click on the control corresponding to the scenario dialogue item (target dialogue item) that he wants to select from the set of scenario selection dialogue items on the display page. For example, the terminal device 110 can receive the target user's selection of the target dialogue item.

[0076] Accordingly, the prompt words input to the first model may be constructed based on the selected target dialogue item to indicate that the first model needs to consider the context information corresponding to the selected target dialogue item when generating the target media content.

[0077] For example, if the target dialogue item indicates a “battle scene”, the prompt word may be used to instruct the first model to generate target media content corresponding to the “battle scene” based on the received media resources.

[0078] In this way, the embodiments of the present disclosure can further increase the user's participation in the generation of media content, thereby enhancing the user's interactive experience.

[0079] In some further embodiments, the terminal device 110 may, for example, obtain output information (eg, media content) of the first model, and may, for example, directly provide the media content as target media content to the target user.

[0080] In some further embodiments, the terminal device 110 and / or the server 130 may further perform corresponding processing on the output information of the first model to generate target media content to be ultimately provided to the target user.

[0081] As an example, the terminal device 110 and / or the server 130 may also perform appropriate editing operations on the media content output by the first model to obtain the final target media content. For example, the terminal device 110 and / or the server 130 may add one or more elements (e.g., text elements, graphic elements, audio elements, etc.) to the media content output by the first model to obtain the final target media content. Alternatively, the terminal device 110 and / or the server 130 may also adjust the size, proportion, and other information of the media content based on the device information of the terminal device 110.

[0082] In some other embodiments, the target media content may also be generated using a second model. It should be understood that such a first model may include, for example, any appropriate machine learning model that can be trained to perform the task of generating media content, and this disclosure is not intended to be limited thereto.

[0083] As an example, the terminal device 110 and / or the server 130 may construct input information for the second model based on at least one dialogue item selected by the user. For example, the terminal device 110 and / or the server 130 may construct a guideline for guiding the second model to generate an image based on the dialogue item selected by the user. For example, if the user selected city A, character B, and weapon C, the guideline could be, for example, "Please generate an image of a hero holding weapon C in a large city."

[0084] Furthermore, the terminal device 110 and / or the server 130 may obtain intermediate media content, such as a sketch, generated by the second model based on the input information. Additionally, the terminal device 110 and / or the server 130 may generate final target media content based on the intermediate media content and the media resources corresponding to the at least one received input item.

[0085] For example, the terminal device 110 and / or the server 130 may replace the character in the sketch with the character selected by the user. For example, the terminal device 110 and / or the server 130 may utilize a machine learning model to combine the intermediate media content with the media resources to provide the final target media content. For example, the target media content may retain the posture of the character in the intermediate media content while being replaced with the character corresponding to the dialogue item selected by the user.

[0086] Based on the process described above, the embodiments of the present disclosure can provide personalized media content based on the user's dialogue interaction in the virtual scene, thereby improving the user's interactive experience in the virtual scene.

[0087] Example devices and equipment

[0088] Embodiments of the present disclosure also provide corresponding apparatuses for implementing the above-described methods or processes. FIG4 shows a schematic structural block diagram of an example apparatus 400 for providing media content according to certain embodiments of the present disclosure. Apparatus 400 may be implemented as or included in terminal device 110. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0089] As shown in Figure 4, the device 400 includes a receiving module 410, which is configured to receive a selection of a set of dialogue items provided in at least one round of dialogue interaction between a target user and a virtual object in a virtual scene; and a providing module 420, which is configured to provide target media content, where the target media content is generated based on a media resource corresponding to at least one dialogue item in the set of dialogue items, and the media resource is associated with the virtual scene.

[0090] In some embodiments, at least one round of conversation includes at least: a first round of conversation interaction and a second round of conversation interaction; apparatus 400 further includes a control module configured to control a virtual object to provide a first statement in the first round of conversation interaction; receiving module 410 is further configured to receive a target user's selection of a first conversation item from a first set of candidate conversation items, the first conversation item indicating a candidate response to the first statement. Furthermore, the control module is further configured to control the virtual object to provide a second statement in the second round of conversation interaction, and receiving module 410 is further configured to receive a second selection of a second conversation item from a second set of candidate conversation items, the second conversation item indicating a candidate response to the second statement by the target user.

[0091] In some embodiments, the second statement and / or the second set of candidate dialogue items are generated based on the selected first dialogue item from the first set of candidate dialogue items.

[0092] In some embodiments, the first group of candidate conversation items or the second group of candidate conversation items includes a plurality of candidate conversation items, and each of the plurality of candidate conversation items corresponds to a different preset media resource.

[0093] In some embodiments, the media resource includes at least one of the following: a first media resource associated with a virtual location included in the virtual scene; a second media resource associated with a virtual prop included in the virtual scene; and a third media resource associated with a virtual character included in the virtual scene.

[0094] In some embodiments, the media resource corresponding to the at least one conversation item includes a target media resource selected from a group of candidate media resources corresponding to the at least one conversation item.

[0095] In some embodiments, the target media resource is selected from a set of candidate media resources based on at least one other conversation item in the set of conversation items.

[0096] In some embodiments, the set of candidate media assets includes a plurality of images associated with the virtual scene, the plurality of images corresponding to different viewpoints.

[0097] In some embodiments, the group of conversation items further includes a target conversation item, and the target media content is further generated based on context information indicated by the target conversation item, where the context information is used to describe a context corresponding to the target media content to be generated.

[0098] In some embodiments, the target media content is generated based on the following process: generating first input information to a first model based on a media resource corresponding to at least one dialog item; and generating target media content generated based on the media resource based on first output information of the first model.

[0099] In some embodiments, the target media content is generated based on the following process: generating second input information to the second model based on at least one dialogue item; obtaining intermediate media content generated by the second model based on the second input information; and generating the target media content based on the intermediate media content and the media resources corresponding to the at least one dialogue item.

[0100] The units included in the device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 400 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0101] FIG5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in FIG5 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. The electronic device 500 shown in FIG5 can be used to implement the terminal device 110 shown in FIG1 .

[0102] As shown in FIG5 , electronic device 500 is a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to enhance the parallel processing capabilities of electronic device 500.

[0103] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the electronic device 500.

[0104] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0105] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0106] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0107] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0108] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0109] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0110] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0111] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0112] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for providing media content, comprising: In at least one round of dialogue interaction between a target user and a virtual object in a virtual scene, receiving a selection of a set of dialogue items provided in the at least one round of dialogue interaction; as well as Target media content is provided, where the target media content is generated based on a media resource corresponding to at least one dialog item in the set of dialog items, where the media resource is associated with the virtual scene.

2. The method according to claim 1, wherein the at least one round of dialogue interaction comprises at least: The first round of dialogue interaction and the second round of dialogue interaction; The first round of dialogue interaction includes: controlling the virtual object to provide a first statement, and receiving a selection by the target user of a first dialogue item in a first group of candidate dialogue items, wherein the first dialogue item indicates a candidate response to the first statement, The second round of dialogue interaction includes: controlling the virtual object to provide a second statement, and receiving a second selection by the target user for a second dialogue item in a second group of candidate dialogue items, wherein the second dialogue item indicates a candidate response to the second statement.

3. The method according to claim 2, wherein the second sentence and / or the second set of candidate dialogue items are generated based on the first dialogue item selected from the first set of candidate dialogue items.

4. The method according to claim 2, wherein the first group of candidate dialogue items or the second group of candidate dialogue items includes multiple candidate dialogue items, and each candidate dialogue item in the multiple candidate dialogue items corresponds to a different preset media resource.

5. The method according to claim 1, wherein the media resource comprises at least one of the following: a first media resource associated with a virtual location included in the virtual scene; a second media resource associated with a virtual prop included in the virtual scene; A third media resource associated with the virtual character included in the virtual scene.

6. The method according to claim 1, wherein the media resource corresponding to the at least one conversation item comprises a target media resource selected from a group of candidate media resources corresponding to the at least one conversation item.

7. The method of claim 6, wherein the target media resource is selected from the set of candidate media resources based on at least one other conversation item in the set of conversation items. 8 . The method according to claim 6 , wherein the set of candidate media resources comprises a plurality of images associated with the virtual scene, the plurality of images corresponding to different viewpoints.

9. The method according to claim 1, wherein the group of dialogue items also includes a target dialogue item, and the target media content is also generated based on the situation information indicated by the target dialogue item, and the situation information is used to describe the situation corresponding to the target media content to be generated.

10. The method according to claim 1, wherein the target media content is generated based on the following process: generating first input information to a first model based on the media resource corresponding to the at least one dialog item; and The target media content generated based on the media resource is generated based on the first output information of the first model.

11. The method according to claim 1, wherein the target media content is generated based on the following process: generating second input information to a second model based on the at least one dialog item; Acquire intermediate media content generated by a second model based on the second input information; as well as The target media content is generated based on the intermediate media content and the media resource corresponding to the at least one dialog item.

12. A device for providing media content, comprising: A selection module is configured to receive, in at least one round of dialogue interaction between a target user and a virtual object in a virtual scene, a selection of a set of dialogue items provided in the at least one round of dialogue interaction; as well as The providing module is configured to provide target media content, wherein the target media content is generated based on a media resource corresponding to at least one dialog item in the group of dialog items, and the media resource is associated with the virtual scene.

13. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 11 when executed by the at least one processing unit.

14. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Interaction method and device, equipment and storage medium

    CN116866402A

  • LNOB video interaction platform

    CN117061806A

  • Human-computer interaction method, device and equipment and storage medium

    CN117111738A

  • Method and device for providing media content, electronic equipment and storage medium

    CN118092731A

  • Media selection and display based on conversation topics

    US20190087498A1