Network live broadcast processing method and device

By generating fusion virtual objects using generative models, the problem of diversity and fun in user interaction during live streaming is solved, thus improving the user's interactive experience in the live streaming room.

CN121284301APending Publication Date: 2026-01-06ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511502317.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-06

AI Technical Summary

Technical Problem

In online live streaming, how can we increase user attention and retention, especially with diverse interaction methods, and how can we enhance the diversity and fun of user interaction in the live stream?

Method used

By receiving multiple virtual objects selected by the user terminal, querying object fusion records, obtaining user feature data, constructing fusion data, and inputting it into a generative model to generate fused virtual objects, which are then displayed and processed for live interactive streaming.

Benefits of technology

It enhances the diversity and fun of virtual object interaction for users in the live streaming room, thus improving the user's live streaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284301A_ABST
    Figure CN121284301A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a network live broadcast processing method and device, and the method comprises the steps: receiving a plurality of virtual objects selected by a user terminal in an interaction list of a live broadcast room in a process that an access user accesses the live broadcast room through the user terminal, and querying an object fusion record of the access user for performing virtual object fusion, if the object fusion record is not queried, constructing fusion data based on the plurality of virtual objects and the user feature data, inputting the fusion data into the generative model for performing fusion virtual object generation to obtain a fusion virtual object, and returning the fusion virtual object to the user terminal. And live broadcast interaction processing of the fused virtual object is carried out, so that live broadcast interaction processing is realized based on the fused virtual object generated by calling the generative model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of data processing technology, and in particular to a method and apparatus for processing live web streaming. Background Technology

[0002] With the continuous development and promotion of the Internet, the application scope of various online services provided by the Internet is becoming wider and wider. In the field of live streaming, users can watch live content and interact in various ways during the process of accessing live streaming, such as sending gifts to the streamer. However, as more and more organizations provide live streaming and the interaction methods of users during the process of accessing live streaming become more and more diversified, how to improve user attention and retention in the live streaming room has become the focus of attention for all parties. Summary of the Invention

[0003] This specification provides one or more embodiments of a live streaming processing method, comprising: receiving multiple virtual objects selected by a user terminal from an interaction list in a live streaming room; querying the object fusion record of the virtual objects fusion performed by the user; if the query result is empty, obtaining user feature data; constructing fusion data based on the multiple virtual objects and the user feature data, and inputting the fusion data into a generative model to generate fused virtual objects; and returning the fused virtual objects to the user terminal for live interactive processing of the fused virtual objects.

[0004] This specification provides one or more embodiments of another online live streaming processing method, including: acquiring multiple virtual objects selected by a user in a live streaming room and submitting them to a server; receiving a fused virtual object returned by the server; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model; and performing display processing on the fused virtual object to enable live interactive processing of the fused virtual object in cooperation with the server.

[0005] This specification provides one or more embodiments of a network live streaming processing apparatus, comprising: an object receiving module configured to receive multiple virtual objects selected by a user terminal from an interaction list in a live streaming room; a record query module configured to query object fusion records of virtual object fusion performed by the user, and if the query result is empty, to obtain user feature data; an object generation module configured to construct fusion data based on the multiple virtual objects and the user feature data, and input the fusion data into a generative model to generate fused virtual objects; and an object return module configured to return the fused virtual objects to the user terminal for live interactive processing of the fused virtual objects.

[0006] This specification provides one or more embodiments of another network live streaming processing apparatus, including: an object submission module configured to acquire multiple virtual objects selected by a user in a live streaming room and submit them to a server; an object receiving module configured to receive a fused virtual object returned by the server; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model; and an object display module configured to perform display processing of the fused virtual object, so as to perform live interactive processing of the fused virtual object in cooperation with the server.

[0007] This specification provides one or more embodiments of a server, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: receive a plurality of virtual objects selected by a user terminal from an interaction list in a live stream; query the object fusion record of the virtual objects being fused by the user, and if the query result is empty, obtain user feature data; construct fusion data based on the plurality of virtual objects and the user feature data, and input the fusion data into a generative model to generate fused virtual objects; and return the fused virtual objects to the user terminal for live interactive processing of the fused virtual objects.

[0008] This specification provides one or more embodiments of a user terminal, including: a processor; and a memory configured to store computer-executable instructions, which, when executed, cause the processor to: acquire and submit multiple virtual objects selected by a user in a live streaming room to a server; receive a fused virtual object returned by the server; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model; and perform display processing of the fused virtual object to conduct live interactive processing of the fused virtual object in cooperation with the server.

[0009] This specification provides one or more embodiments of a computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: receiving multiple virtual objects selected by a user terminal from an interaction list in a live broadcast room; querying the object fusion record of the virtual object fusion performed by the user; if the query result is empty, obtaining user feature data; constructing fusion data based on the multiple virtual objects and the user feature data, and inputting the fusion data into a generative model to generate fused virtual objects; and returning the fused virtual objects to the user terminal for live broadcast interaction processing of the fused virtual objects.

[0010] This specification provides one or more embodiments of another computer-readable storage medium for storing computer-executable instructions, which, when executed, implement the following process: acquiring multiple virtual objects selected by a user in a live stream and submitting them to a server; receiving a fused virtual object returned by the server; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model; and performing display processing of the fused virtual object to enable live interactive processing of the fused virtual object in cooperation with the server. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 A schematic diagram illustrating the implementation environment of a live streaming processing method provided in one or more embodiments of this specification; Figure 2 A flowchart illustrating a network live streaming processing method provided in one or more embodiments of this specification; Figure 3 A schematic diagram of a first type of live streaming page provided for one or more embodiments of this specification; Figure 4 A schematic diagram of a second type of live streaming page provided for one or more embodiments of this specification; Figure 5 A schematic diagram of a third type of live streaming page provided for one or more embodiments of this specification; Figure 6 A timing diagram of a network live streaming processing method applied to a live streaming platform scenario, provided by one or more embodiments of this specification; Figure 7 A flowchart of another network live streaming processing method provided in one or more embodiments of this specification; Figure 8 A schematic diagram of an embodiment of a network live streaming processing device provided in one or more embodiments of this specification; Figure 9 A schematic diagram of another embodiment of a network live streaming processing apparatus provided in one or more embodiments of this specification; Figure 10 A schematic diagram of the structure of a server provided for one or more embodiments of this specification; Figure 11This is a schematic diagram of the structure of a user terminal provided for one or more embodiments of this specification. Detailed Implementation

[0012] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0013] The online live streaming processing method provided in one or more embodiments of this specification is applicable to the implementation environment of an online live streaming platform. (Refer to...) Figure 1 The implementation environment includes at least: Accessing the user's user terminal 101, and providing the live streaming room server 102; Among them, the user terminal 101 of the accessing user is used to access the live broadcast room and process live broadcast interaction, and cooperates with the server 102 to perform corresponding processing during the live broadcast process; the user terminal 101 of the accessing user can be a mobile phone, personal computer, tablet computer, e-book reader, VR (Virtual Reality) based information interaction device, vehicle terminal, IoT device, wearable smart device, laptop computer and desktop computer, etc. Server 102 is used to respond to and process the live interaction of the user terminal 101 during the live broadcast. Specifically, server 102 can receive multiple virtual objects selected by user terminal 101 in the interaction list of the live broadcast room, generate a fused virtual object based on multiple virtual objects and user feature data, and return the fused virtual object to user terminal 101. Server 102 may be deployed with a generative model 102-1, which may be configured with corresponding modules or networks to perform related processing for virtual object generation. Server 102 may be a single server, a server cluster consisting of several servers, or one or more cloud servers in a cloud computing platform.

[0014] The implementation environment may also include the live streaming terminal 103 of the broadcaster; the live streaming terminal 103 of the broadcaster can be used to access the live streaming room and process live streaming interaction, and cooperate with the server 102 to perform corresponding processing during the live streaming process; the live streaming terminal 103 of the broadcaster may specifically be a mobile phone, personal computer, tablet computer, e-book reader, VR (Virtual Reality) based information interaction device, vehicle terminal, IoT device, wearable smart device, laptop computer and desktop computer, etc.

[0015] In this implementation environment, during the process of a user accessing a live broadcast room through a user terminal, server 102 receives multiple virtual objects selected by user terminal 101 from the interaction list of the live broadcast room, and queries the object fusion record of the virtual objects fusion performed by the user. If server 102 does not find an object fusion record, it constructs fusion data based on multiple virtual objects and user feature data, and inputs the fusion data into a generative model to generate fused virtual objects. After that, server 102 returns the fused virtual object to user terminal 101. User terminal 101 receives the fused virtual object returned by server 102 and performs the display processing of the fused virtual object. Based on this, user terminal 101 performs live interactive processing of the fused virtual object in cooperation with server 102, thereby realizing live interactive processing by calling the fused virtual object generated by the generative model through server.

[0016] It should be noted that, considering that the user characteristic data and other related data involved in this manual may, to some extent, be considered user privacy, user authorization should be obtained before collecting such data to ensure that the data collection operation complies with relevant data management regulations. For example, data authorization can be granted by sending a data authorization reminder to the user and having the user confirm the data authorization.

[0017] This specification provides one or more embodiments of a network live streaming processing method as follows: Reference Figure 2 The live streaming processing method provided in this embodiment can be applied to a server, and the method specifically includes steps S202 to S208.

[0018] Step S202: Receive multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room.

[0019] The virtual object described in this embodiment refers to a virtualized digital object constructed based on data elements and associated with a live streaming scenario, possessing certain resource-related or permission-related attributes. Specifically, a virtual object can be a virtual interactive object that a user interacts with the streamer during their visit to the live streaming room, such as a virtual interactive object gifted by a user to the streamer. Furthermore, a virtual object can also be a virtual interactive object that a user interacts with during their visit to the live streaming room, or with the live streaming provider or service provider; alternatively, it can be a virtual interactive object that a user interacts with during their visit to the live streaming room, or with tasks or associated projects configured within the live streaming room.

[0020] An interactive list refers to an interactive component used to provide users with opportunities to interact with virtual objects. For example, an interactive list can be a list of virtual objects configured in multiple live streaming rooms, or it can be an interactive component configured with multiple virtual objects.

[0021] In practice, when a user visits a live stream, they interact with the live stream by selecting multiple virtual objects from the interaction list. The user's terminal uploads these virtual objects to the server. Correspondingly, the server receives the virtual objects selected by the user's terminal from the interaction list, which means it receives the virtual objects selected by the user's terminal from the interaction list.

[0022] In this embodiment, based on multiple virtual objects selected from the interaction list in the live broadcast room, fused data is constructed by combining relevant data from the selected virtual objects. The fused data is then input into a generative model to generate fused virtual objects, thereby providing fused virtual objects for live broadcast interaction to the visiting users. In this scenario, to improve the visiting users' perception of the fused virtual objects, before generating the fused virtual objects, the visiting users can be shown the fused object identifiers corresponding to multiple virtual objects, so that the visiting users can perceive and generate fused virtual objects that meet their expectations through the fused object identifiers.

[0023] In one optional implementation of this embodiment, after receiving multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room, the following operations are performed: Read the fusion object identifiers corresponding to multiple virtual objects and return them to the user terminal; If a fusion confirmation instruction is received from the user terminal, the following query operation is performed to access the object fusion record of the user performing virtual object fusion; Optionally, the fusion object identifier is obtained by inputting multiple virtual objects and fusion task templates into the generative model to generate fusion virtual object identifiers; in addition, the fusion object identifier can also be the fusion object identifier of the fusion virtual object generated after the historical access user selects multiple virtual objects stored on the server, or it can be the fusion object identifier generated based on the fusion virtual object identifier of the historical access user.

[0024] In addition, after reading the fusion object identifiers corresponding to multiple virtual objects and returning them to the user terminal, if a switching instruction submitted by the user terminal is received, an updated object identifier corresponding to multiple virtual objects is generated and returned to the user terminal. Specifically, in the process of generating the updated object identifiers corresponding to multiple virtual objects, the updated object identifier can be obtained by adjusting the fusion object identifier, or by adjusting the object relationship of multiple virtual objects and inputting the adjusted multiple virtual objects and the fusion task template into the generative model to generate the fusion virtual object identifier. Alternatively, the merged object identifier of the merged virtual object generated after the historical users selected multiple virtual objects stored on the server can be input into the generative model for object identifier adjustment; or, the merged object identifier generated based on the object identifier of the merged virtual object of the historical users can be input into the generative model for object identifier adjustment.

[0025] Step S204: Query the object fusion records of the user performing virtual object fusion. If the query result is empty, obtain the user feature data.

[0026] In practice, after receiving multiple virtual objects selected by the user terminal in the interactive list of the live broadcast room, the system queries the object fusion record of the virtual object fusion performed by the user. If the query result is empty, it indicates that the current user has not previously selected multiple virtual objects for virtual object fusion in the live broadcast room. Then, the user feature data is obtained, and the corresponding virtual object fusion is performed based on the multiple virtual objects selected by the user and the user feature data. If the query result is not empty, that is, if the object fusion record of the visiting user is found, it means that the current visiting user has previously selected multiple virtual objects in the live broadcast room for virtual object fusion. Then, the corresponding virtual object fusion will be performed based on the multiple virtual objects selected by the visiting user and the object fusion record.

[0027] The fused virtual object described in this embodiment is a new virtual object generated by analyzing the resource amount relationship and the relationship between the virtual objects, and combining the fusion task template, user interaction features and / or object fusion records, based on multiple virtual objects. It is not a fusion result obtained by directly fusing the object data of multiple virtual objects.

[0028] In practice, based on the retrieved object fusion records, fusion data can be constructed using multiple virtual objects and these records. This fusion data is then input into a generative model to generate fused virtual objects. Subsequently, the fused virtual objects can be returned to the user terminal for live interactive processing, thereby enhancing the diversity and engagement of virtual object interactions for users in the live stream.

[0029] Specifically, in the process of generating fused virtual objects, feature extraction can be performed on the object fusion record to obtain record features, and image feature extraction can be performed on multiple virtual objects to obtain object image features. Cross-modal feature fusion can be performed on the record features and object image features, and diffusion and feature decoding can be performed based on the fused features to obtain fused virtual objects.

[0030] Furthermore, to enhance the live stream interaction experience for users by selecting multiple virtual objects in the live stream room and then merging them into a single virtual object, the generation process can incorporate festival characteristics corresponding to festival information. This allows the generated virtual objects to reflect festival elements, improving the user interaction in the live stream room. Specifically, in one optional implementation of this embodiment, the generation of merged virtual objects includes: The text features of the object fusion record are mapped with the image features of multiple virtual objects, and the fused features are generated based on the feature mapping results. The fusion object features are generated based on the fusion features and the fusion sub-features extracted from the text features and image features. The fusion object features are then transformed into pixels to obtain the fusion object image.

[0031] In the specific execution process, firstly, text features are extracted from the object fusion record to obtain text features, and then image features are extracted from multiple virtual objects to obtain multiple image features. Then, the text features of the object fusion record are mapped with the image features of multiple virtual objects. Based on the feature mapping results, object fusion rules for fusing multiple virtual objects are generated, and fusion features are generated based on the object fusion rules. Furthermore, based on the object fusion rules, sub-features are extracted from text features and multiple image features, and the extracted sub-features are concatenated into fused object features. Alternatively, the extracted sub-features are concatenated into initial fused features, and feature supplementation and feature reconstruction are performed on the initial fused features to obtain fused object features. Finally, pixel transformation is performed on the fused object features to obtain a fused object image as a fused virtual object.

[0032] It should be noted that, in the process of generating fused virtual objects based on multiple virtual objects and object fusion records, in addition to the above-mentioned implementation method for generating fused virtual objects, a similar implementation method can be used as described below. Specifically, the user feature data in the process described below can be replaced with object fusion records to form a new implementation method for generating fused virtual objects based on multiple virtual objects and object fusion records. Alternatively, the implementation method described below can be used to replace or modify the implementation method for generating fused virtual objects based on multiple virtual objects and object fusion records to obtain a new implementation method. Specifically, any one or more of the following implementation methods for constructing fused data based on multiple virtual objects and user feature data and inputting the fused data into a generative model to generate fused virtual objects can be combined with the implementation method here for constructing fused data based on multiple virtual objects and object fusion records and inputting the fused data into a generative model to generate fused virtual objects to create a new implementation method; In addition, any one or more of the following implementation methods for constructing fused data based on multiple virtual objects and user feature data and inputting the fused data into a generative model to generate fused virtual objects can be used to replace one or more of the implementation methods for constructing fused data based on multiple virtual objects and object fusion records and inputting the fused data into a generative model to generate fused virtual objects to form a new implementation method.

[0033] Step S206: Construct fused data based on the multiple virtual objects and the user feature data, and input the fused data into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects.

[0034] In practice, if the query result for the object fusion record of the user's virtual object fusion is empty, meaning the current user has not previously selected multiple virtual objects for virtual object fusion in the live stream, then user feature data is obtained. Based on the multiple virtual objects selected by the user and combined with the user feature data, a fused virtual object is generated. Specifically, fusion data is constructed based on multiple virtual objects and user feature data, and this fusion data is input into a generative model to generate the fused virtual object. The generative model refers to a large language model capable of generating virtual objects, and this large language model is configured with corresponding modules or networks capable of performing related processing for virtual object generation.

[0035] To improve the generation effect of fused virtual objects and make the generated fused virtual objects match the intentions and needs of users who are currently interacting with the live broadcast in the live broadcast room as much as possible, in the process of building fused data and inputting the fused data into the generative model to generate fused virtual objects, it is also possible to combine the pre-configured fused task template for virtual object fusion, and generate fused virtual objects from three aspects of data: multiple virtual objects, user feature data, and fused task template.

[0036] Specifically, in one optional implementation of this embodiment, fused data is constructed based on multiple virtual objects and user feature data, including: writing user feature data into a fused task template to obtain task data, and associating multiple virtual objects with the task data to obtain fused data.

[0037] Optionally, the fusion task template may contain at least one of the following: a subject role field for fusion of virtual objects, an input rule field, a generation task field for generating fusion virtual objects, and an output rule field.

[0038] In the specific execution process, during the process of generating fused virtual objects based on the input fused data, the generative model can perform feature mapping between the task features of the task data and the image features of multiple virtual objects, generate fused object features based on the feature mapping results, task features and image features, and perform pixel transformation on the fused object features to obtain fused virtual objects. Specifically, in one optional implementation of this embodiment, the generation of fused virtual objects includes: The task features of the task data are mapped to the image features of multiple virtual objects, and fused features are generated based on the feature mapping results. Based on the fusion features and the fusion sub-features extracted from the task features and image features, fusion object features are generated, and pixel transformation is performed on the fusion object features to obtain a fusion virtual object.

[0039] In the specific execution process, firstly, feature extraction is performed on the task data to obtain task features, and image feature extraction is performed on multiple virtual objects to obtain multiple image features. Then, feature mapping is performed between the task features and the image features of multiple virtual objects. Based on the feature mapping results, object fusion rules for fusing multiple virtual objects are generated, and fusion features are generated based on the object fusion rules. Further, sub-features are extracted from text features and multiple image features according to the object fusion rules, and the extracted sub-features are concatenated to form fusion object features. Alternatively, the extracted sub-features are concatenated to form initial fusion features, and feature supplementation and feature reconstruction are performed on the initial fusion features to obtain fusion object features. Finally, pixel conversion is performed on the fusion object features to obtain the fusion object image as the fusion virtual object.

[0040] In addition, during the process of generating fusion features based on feature mapping results, the object fusion rules corresponding to the fusion features may also include the main virtual object and subordinate virtual objects among multiple virtual objects; Specifically, the primary and subordinate virtual objects among multiple virtual objects are determined in the following ways: the resource amount relationship between the resource amounts corresponding to the multiple virtual objects is determined, and the primary and subordinate virtual objects are determined based on the resource amount relationship and the resource characteristics contained in the theme role characteristics; or, the primary and subordinate virtual objects are determined based on the resource amounts corresponding to each of the multiple virtual objects.

[0041] The resource amount refers to the amount of resources that users need to pay to interact with the streamer, the live stream provider, or the live stream service provider by selecting virtual objects. The resources can be virtual resources or funds, such as carbon savings.

[0042] A primary virtual object refers to a virtual object that serves as the main reference or plays a major role when a virtual object is generated by merging multiple virtual objects; a subordinate virtual object refers to a virtual object that serves as a secondary reference or plays a minor role when a virtual object is generated by merging multiple virtual objects.

[0043] During the process of generating fused virtual objects, decisions can also be made regarding the resource amount of the fused virtual objects. For example, after obtaining fused virtual objects by pixel conversion of the features of the fused objects, the resource amount corresponding to the fused virtual objects can be determined based on the resource amount of each of the multiple virtual objects and the theme role field.

[0044] Specifically, in one optional implementation of this embodiment, the generation of fused virtual objects further includes: Extract the object resource features corresponding to the resource amounts of multiple virtual objects, and extract the resource features of the theme role field to obtain the theme resource features; Perform feature matching between object resource features and topic resource features, and determine resource weights based on the feature matching results; The resource amount corresponding to the merged virtual object is determined by the resource weight and the resource amount corresponding to each of the multiple virtual objects.

[0045] In the process of determining the resource amount of the merged virtual object based on the resource amount of each of the multiple virtual objects and the theme role field, the object resource features of the resource amount of each of the multiple virtual objects are first extracted, and the theme resource features are obtained by extracting the resource features of the theme role field. Then, the object resource features and theme resource features are matched, and the resource weight is determined based on the feature matching results. Further, the resource amount is calculated based on the resource weight and the resource amount of each of the multiple virtual objects to obtain the initial resource amount, and the initial resource amount is adjusted based on the theme resource features to obtain the resource amount of the merged virtual object.

[0046] In addition, in determining the resource amount corresponding to the merged virtual object, the resource amount can also be determined based on the resource amount of each virtual object and its relationship characteristics with the theme, master-slave relationship, and / or user characteristics. Specifically, determining the resource amount based on the resource amount of each virtual object and its relationship characteristics with the theme, master-slave relationship, and / or user characteristics is similar to the implementation method described above. User characteristics refer to the features obtained by feature extraction from user feature data.

[0047] In practical applications, during the generation of fused virtual objects, the generative model can further generate fused virtual objects that better match the relationships between the multiple virtual objects by using the master and subordinate virtual objects among the multiple virtual objects. Specifically, in one optional implementation provided in this embodiment, the generation of fused virtual objects by the generative model includes: The main virtual object and subordinate virtual objects are determined based on the image semantic features of multiple virtual objects, or based on the object type and / or user characteristics of multiple virtual objects. The main object features of the main virtual object and the subordinate object features of the subordinate virtual objects are fused to obtain fused features. Based on the task features of the task data, the fused features are adjusted and pixel transformed to obtain the fused virtual object.

[0048] Furthermore, during the generation of fused virtual objects, the generative model can also generate fused virtual objects that better match the relationships between the multiple virtual objects by using the master virtual object and subordinate virtual objects among the multiple virtual objects. Specifically, in another optional implementation provided in this embodiment, the generation of fused virtual objects by the generative model includes: User features and holiday features are mapped to image features of multiple virtual objects, and fused features are generated based on the feature mapping results. The fusion object features are generated based on the fusion features and the fusion sub-features extracted from the festival features and image features. Then, pixel transformation is performed on the fusion object features to obtain the fused virtual object. Here, the festival features refer to the features obtained by extracting features from the festival information of the current festival.

[0049] It should be noted that the various implementation methods provided above for constructing fused data based on multiple virtual objects and user feature data, and inputting the fused data into a generative model to generate fused virtual objects, can be combined in any form to form new implementation methods according to actual execution needs. Alternatively, one or more steps from different implementation methods can be combined to form new implementation methods suitable for actual execution scenarios. For example, generating fused virtual objects includes: determining the main virtual object and subordinate virtual objects based on the image semantic features, object types, and / or user features of multiple virtual objects; performing feature fusion on the main object features of the main virtual object and the subordinate object features of the subordinate virtual objects to obtain fused features, and performing feature adjustment and pixel transformation on the fused features based on the task features of the task data to obtain fused virtual objects; performing feature matching between the object resource features corresponding to the resource amounts of each of the multiple virtual objects and the theme resource features of the theme role field (or the task features of the task data), and determining the resource weights based on the feature matching results; and determining the resource amounts corresponding to the fused virtual objects based on the resource weights and the resource amounts corresponding to each of the multiple virtual objects.

[0050] Step S208: Return the fused virtual object to the user terminal to perform live interactive processing of the fused virtual object.

[0051] Based on the above-mentioned generation of fused virtual objects through a generative model, the obtained fused virtual objects are then returned to the user terminal for live interactive processing. Specifically, the live interactive processing of fused virtual objects can involve the user confirming the interaction with the fused virtual objects, and the user terminal sending an interaction confirmation instruction to the server based on the interaction confirmation. Correspondingly, after receiving the interaction confirmation instruction, the server deducts resources from the user's resource account to perform the interactive processing of the fused virtual objects.

[0052] For example, Figure 3 The interactive list shown is the live page of the live room accessed by the user through the user terminal. The interactive list consists of multiple virtual objects. The user selects virtual object 301 and virtual object 302 in the interactive list and initiates the fusion of virtual objects by triggering the "Merge" button configured on the live page. After the user terminal submits the virtual object 301 and virtual object 302 selected by the user in the interactive list to the server, the server queries the object fusion record of the virtual object fusion performed by the user. If the query result is empty, the user feature data is obtained. Based on virtual object 301, virtual object 302, and user feature data, fused data is constructed. This fused data is then input into a generative model to generate fused virtual objects. The generated fused virtual objects are returned to the user terminal, and the fused virtual objects displayed on the user terminal are as follows: Figure 4 As shown.

[0053] Furthermore, after returning the obtained fused virtual object to the user terminal, the live interaction processing of the user regarding the fused virtual object can also involve updating the fused virtual object. Specifically, in one optional implementation of this embodiment, the live interaction processing of the fused virtual object includes: If an object update request is received from a user terminal, the object update data corresponding to the object update request is generated and input into the generative model to adjust the fused virtual object, thereby obtaining the updated virtual object; Alternatively, if an object update request is received from a user terminal, update fusion data is constructed based on the object update request, multiple virtual objects, and user feature data. The update fusion data is then input into a generative model to generate fused virtual objects, thereby obtaining updated virtual objects.

[0054] Specifically, in one optional implementation of this embodiment, adjustments are made to the merged virtual object, including: Adjust the master-slave relationship between the master virtual object and the subordinate virtual object among the multiple virtual objects stored in the generative model, and extract the adjusted master object features and subordinate object features; Virtual objects are generated based on the characteristics of the subject role, user, main object, and subordinate object, and updated virtual objects are obtained.

[0055] Following the previous example, the user terminal is displayed. Figure 4 Based on the merged virtual object shown, if a user updates the merged virtual object by clicking the "Change" button, the updated virtual object returned by the server after the update is as follows: Figure 5 As shown; where, Figure 4 It is a fused virtual object generated by fusing virtual object 301 as the main virtual object and virtual object 302 as the subordinate virtual object; Figure 5 It is a fused virtual object generated by the server after adjusting the master-slave relationship between virtual objects 301 and virtual objects 302, with virtual object 302 as the master virtual object and virtual object 301 as the subordinate virtual object.

[0056] It should be noted that the process of constructing updated fusion data based on object update requests, multiple virtual objects, and user feature data, and then inputting the fusion data into the generative model to generate fused virtual objects, can be implemented in a similar way to the process described above of constructing fusion data based on multiple virtual objects and user feature data and then inputting the fusion data into the generative model to generate fused virtual objects. For example, adjusting the fused virtual object includes: performing feature mapping between the update features of the object update request and the fused object features of the fused virtual object, generating update features based on the feature mapping results; generating update object features based on the update features and the update sub-features extracted from the fused object features, and performing pixel transformation on the update object features to obtain the updated virtual object; For example, adjusting the fused virtual object includes: performing feature mapping between user features and festival features and image features of multiple virtual objects, generating updated features based on the feature mapping results; generating updated object features based on the updated features and the fused sub-features extracted from the festival features and image features, and performing pixel transformation on the updated object features to obtain the updated virtual object; Alternatively, feature mapping can be performed between the festival features and the fusion object features of the fused virtual object, and updated features can be generated based on the feature mapping results; updated object features can be generated based on the updated features and the updated sub-features extracted from the festival features and / or fused object features, and the updated object features can be pixel-transformed to obtain the updated virtual object.

[0057] In this embodiment, when multiple virtual objects are selected by the user terminal in the interaction list of the live room, the live room can be checked for anomalies. If the check passes, asynchronous interaction tasks corresponding to multiple virtual objects are created and added to the asynchronous task queue. In this way, the asynchronous task queue can realize parallel virtual object interaction, that is, the generative model can be called in parallel to generate fused virtual objects, which helps to improve the response efficiency and real-time interaction of virtual object interaction in the live room.

[0058] Subsequently, during the processing of asynchronous interactive tasks in the asynchronous task queue, based on the generation of a fusion object or the updating of a virtual object, the system checks whether the interaction status of the fusion object or the updating of the virtual object is normal. If so, it reads the resource amount of the fusion object or the updating of the virtual object and checks whether the resource amount in the accessing user's resource account is greater than or equal to the resource amount of the fusion object or the updating of the virtual object. If it is greater, it creates an object interaction record for the accessing user and configures the order status of the object interaction order in the object interaction record to the initialization state. Furthermore, it calls the resource system corresponding to the accessing user's resource account to deduct resources. If the deduction is successful, it updates the order status of the object interaction order in the object interaction record to the paid state. If the deduction fails, it cancels the object interaction order and updates the order status of the object interaction order to the canceled state.

[0059] In summary, the live streaming processing method provided in this embodiment receives multiple virtual objects selected by the user terminal from the interaction list in the live streaming room during the process of the user accessing the live streaming room through the user terminal. It queries the object fusion record of the virtual object fusion performed by the user. Based on the fused object record, it can construct fusion data based on multiple virtual objects and the object fusion record, and input the fusion data into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects that conform to the user's object fusion habits and preferences. Subsequently, it can also return the fused virtual objects to the user terminal for live interactive processing of fused virtual objects, thereby improving the interactive diversity of the user's virtual object interaction in the live streaming room. If no object fusion record is found, fusion data is constructed based on multiple virtual objects and user feature data. The fusion data is then input into a generative model to generate fusion virtual objects. The fusion virtual objects are then returned to the user terminal for live interactive processing, thereby enhancing the interactive fun of users interacting with virtual objects in the live broadcast room.

[0060] The following example uses a live streaming processing method provided in this embodiment in a live streaming platform scenario as an example, combined with... Figure 6The network live streaming processing method provided in this embodiment will be further explained below. See [link to documentation]. Figure 6 The network live streaming processing method applied to live streaming platform scenarios includes the following steps.

[0061] Step S604: Receive multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room.

[0062] Step S606: Query and access the object fusion records of the user performing virtual object fusion.

[0063] Step S608: If the query result is empty, obtain user feature data.

[0064] Step S610: Write user feature data into the fusion task template to obtain task data, and associate multiple virtual objects with the task data to obtain fusion data.

[0065] Step S612: Input the fused data into the generative model to generate fused virtual objects and obtain fused virtual objects.

[0066] Step S614: Return the merged virtual object to the user terminal.

[0067] Step S622: Receive an object update request submitted by the user terminal.

[0068] Step S624: Generate object update data corresponding to the object update request and input it into the generative model to adjust the fused virtual object to obtain the updated virtual object.

[0069] Step S626: Return the updated virtual object to the user terminal.

[0070] It should be noted that any one or more of steps S604 to S614 and steps S622 to S626 can be combined with any one or more of steps S202 to S208 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S604 to S614 and steps S622 to S626 can be selected and combined with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S604 to S614 and steps S622 to S626 can also be replaced with any one or more technical features provided in steps S202 to S208 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.

[0071] Furthermore, it should be noted that steps S604 to S614 and steps S622 to S626 provided in this embodiment can be executed by the server. It should be noted that the steps S604 to S614 and steps S622 to S626 executed by the server can cooperate with steps S602, steps S616 to S620 and steps S628 executed by the user terminal in the following embodiment during execution. Therefore, when reading this embodiment, please refer to the corresponding content of steps S602, steps S616 to S620 and steps S628 provided in the following method embodiment, and when reading the following method embodiment, please refer to the corresponding content of steps S604 to S614 and steps S622 to S626 provided in this embodiment.

[0072] One or more embodiments of another online live streaming processing method provided in this specification are as follows: Reference Figure 7 The live streaming processing method provided in this embodiment can be applied to user terminals. The method specifically includes steps S702 to S706.

[0073] Step S702: Obtain the multiple virtual objects selected by the user in the live broadcast room and submit them to the server.

[0074] The virtual object described in this embodiment refers to a virtualized digital object constructed based on data elements and associated with a live streaming scenario, possessing certain resource-related or permission-related attributes. Specifically, a virtual object can be a virtual interactive object that a user interacts with the streamer during their visit to the live streaming room, such as a virtual interactive object gifted by a user to the streamer. Furthermore, a virtual object can also be a virtual interactive object that a user interacts with during their visit to the live streaming room, or with the live streaming provider or service provider; alternatively, it can be a virtual interactive object that a user interacts with during their visit to the live streaming room, or with tasks or associated projects configured within the live streaming room.

[0075] Specifically, users can select multiple virtual objects from the interaction list in the live stream. The interaction list refers to the interactive component used to provide users with the means to interact with virtual objects. For example, the interaction list can be a list of virtual objects configured in multiple live streams, or it can be an interaction component configured with multiple virtual objects.

[0076] In practice, when a user visits a live stream, they interact with the live stream by selecting multiple virtual objects from the interaction list. The user's terminal uploads these virtual objects to the server. The server then receives the virtual objects selected by the user in the interaction list, which is equivalent to receiving the virtual objects selected by the user in the interaction list submitted by the user.

[0077] In this embodiment, the server selects multiple virtual objects from the interactive list in the live broadcast room, and constructs fused data based on the selected virtual objects and related data. The fused data is then input into a generative model to generate fused virtual objects, thereby providing fused virtual objects for live broadcast interaction to the visiting user. In this scenario, to improve the visiting user's perception of the fused virtual objects, the server can also display the fused object identifiers corresponding to the multiple virtual objects to the visiting user before generating the fused virtual objects, so that the visiting user can perceive and generate the expected fused virtual objects through the fused object identifiers.

[0078] In one optional implementation of this embodiment, after receiving multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room, the server performs the following operations: Read the fusion object identifiers corresponding to multiple virtual objects and return them to the user terminal; If a fusion confirmation instruction is received from the user terminal, the following query operation is performed to access the object fusion record of the user performing virtual object fusion; Optionally, the fusion object identifier is obtained by inputting multiple virtual objects and fusion task templates into the generative model to generate fusion virtual object identifiers; in addition, the fusion object identifier can also be the fusion object identifier of the fusion virtual object generated after the historical access user selects multiple virtual objects stored on the server, or it can be the fusion object identifier generated based on the fusion virtual object identifier of the historical access user.

[0079] In addition, after the server reads the fusion object identifiers corresponding to multiple virtual objects and returns them to the user terminal, if a switching instruction is submitted to the server, that is, if the server receives the switching instruction submitted by the user terminal, it generates updated object identifiers corresponding to multiple virtual objects and returns them to the user terminal. Specifically, in the process of generating updated object identifiers corresponding to multiple virtual objects, the updated object identifiers can be obtained by adjusting the fusion object identifiers, or by adjusting the object relationships of multiple virtual objects and inputting the adjusted multiple virtual objects and the fusion task template into the generative model to generate fusion virtual object identifiers. Alternatively, the merged object identifier of the merged virtual object generated after the historical users selected multiple virtual objects stored on the server can be input into the generative model for object identifier adjustment; or, the merged object identifier generated based on the object identifier of the merged virtual object of the historical users can be input into the generative model for object identifier adjustment.

[0080] In practice, after the server receives multiple virtual objects selected by the user terminal in the interactive list of the live broadcast room, it queries the object fusion record of the virtual object fusion performed by the user. If the query result is empty, it means that the current user has not previously selected multiple virtual objects for virtual object fusion in the live broadcast room. Then the server obtains user feature data and uses it to perform corresponding virtual object fusion based on the multiple virtual objects selected by the user and the user feature data. If the query result is not empty, that is, if the server finds the object fusion record of the visiting user, it means that the current visiting user has previously selected multiple virtual objects in the live broadcast room for virtual object fusion. Then, the corresponding virtual object fusion will be performed based on the multiple virtual objects selected by the visiting user and the object fusion record.

[0081] The fused virtual object described in this embodiment is a new virtual object generated by analyzing the resource amount relationship and the relationship between the virtual objects, and combining the fusion task template, user interaction features and / or object fusion records, based on multiple virtual objects. It is not a fusion result obtained by directly fusing the object data of multiple virtual objects.

[0082] In the specific execution process, based on the object fusion records retrieved by the server, fusion data can be constructed based on multiple virtual objects and object fusion records. This fusion data is then input into a generative model to generate fused virtual objects. Subsequently, the server can return the fused virtual objects to the user terminal. Correspondingly, this section receives and displays the returned fused virtual objects to enable live interactive processing of the fused virtual objects, thereby enhancing the diversity and engagement of virtual object interaction for users in the live stream.

[0083] Specifically, during the process of generating fused virtual objects, the server can extract features from the object fusion record to obtain record features, extract image features from multiple virtual objects, perform cross-modal feature fusion of record features and object image features, and perform diffusion and feature decoding based on the fused features to obtain fused virtual objects.

[0084] Furthermore, to enhance the live stream interaction experience for users who select multiple virtual objects in the live stream room to generate a merged virtual object, the server can also incorporate festival characteristics corresponding to festival information into the merged virtual object generation process. This allows the generated merged virtual object to reflect festival elements, improving the interactive experience for users in the live stream room. Specifically, in one optional implementation of this embodiment, the generation of merged virtual objects includes: The text features of the object fusion record are mapped with the image features of multiple virtual objects, and the fused features are generated based on the feature mapping results. The fusion object features are generated based on the fusion features and the fusion sub-features extracted from the text features and image features. The fusion object features are then transformed into pixels to obtain the fusion object image.

[0085] In the specific execution process, firstly, the server extracts text features from the object fusion record to obtain text features, and then extracts image features from multiple virtual objects to obtain multiple image features. Then, the server performs feature mapping between the text features of the object fusion record and the image features of multiple virtual objects, generates object fusion rules for fusing multiple virtual objects based on the feature mapping results, and generates fusion features based on the object fusion rules; Furthermore, the server extracts sub-features from text features and multiple image features according to the object fusion rules, and concatenates the extracted sub-features into fused object features, or concatenates the extracted sub-features into initial fused features and performs feature supplementation and feature reconstruction on the initial fused features to obtain fused object features; finally, the server performs pixel conversion on the fused object features to obtain a fused object image as a fused virtual object.

[0086] It should be noted that, in the process of generating fused virtual objects by constructing fused data based on multiple virtual objects and object fusion records and inputting the fused data into a generative model, in addition to the server-based fused virtual object generation method provided above, a similar implementation method can be adopted as described below. Specifically, the user feature data in the process described below can be replaced with object fusion records to form a new implementation method. Alternatively, the implementation method described below can be used to replace or modify the implementation method described here. Specifically, any one or more of the following implementation methods for constructing fused data based on multiple virtual objects and user feature data and inputting the fused data into a generative model to generate fused virtual objects can be combined with the implementation method here for constructing fused data based on multiple virtual objects and object fusion records and inputting the fused data into a generative model to generate fused virtual objects to create a new implementation method; In addition, any one or more of the following implementation methods for constructing fused data based on multiple virtual objects and user feature data and inputting the fused data into a generative model to generate fused virtual objects can be used to replace one or more of the implementation methods for constructing fused data based on multiple virtual objects and object fusion records and inputting the fused data into a generative model to generate fused virtual objects to form a new implementation method.

[0087] Step S704: Receive the fused virtual object returned by the server.

[0088] In practice, if the query result for the virtual object fusion record of the user is empty, meaning that the current user has not previously selected multiple virtual objects for virtual object fusion in the live broadcast room, the server obtains the user's feature data. Based on the multiple virtual objects selected by the user, the server combines the user's feature data to generate fused virtual objects. Specifically, fusion data is constructed based on multiple virtual objects and user feature data, and the fusion data is input into a generative model to generate fused virtual objects, thus obtaining fused virtual objects.

[0089] Optionally, fused virtual objects are obtained by inputting fused data, constructed based on multiple virtual objects and user feature data, into a generative model to generate fused virtual objects.

[0090] The generative model refers to a large language model capable of generating virtual objects, and this large language model is configured with corresponding modules or networks capable of performing related processing for virtual object generation.

[0091] To improve the generation effect of fused virtual objects and make the generated fused virtual objects match the intentions and needs of users currently interacting in the live broadcast room as much as possible, the server can also combine pre-configured fused task templates for virtual object fusion to generate fused virtual objects from three aspects: multiple virtual objects, user feature data, and fused task templates.

[0092] Specifically, in one optional implementation of this embodiment, fused data is constructed based on multiple virtual objects and user feature data, including: writing user feature data into a fused task template to obtain task data, and associating multiple virtual objects with the task data to obtain fused data.

[0093] Optionally, the fusion task template may contain at least one of the following: a subject role field for fusion of virtual objects, an input rule field, a generation task field for generating fusion virtual objects, and an output rule field.

[0094] In the specific execution process, during the process of generating fused virtual objects based on the input fused data, the generative model can perform feature mapping between the task features of the task data and the image features of multiple virtual objects, generate fused object features based on the feature mapping results, task features and image features, and perform pixel transformation on the fused object features to obtain fused virtual objects. Specifically, in one optional implementation of this embodiment, the generation of fused virtual objects includes: The task features of the task data are mapped to the image features of multiple virtual objects, and fused features are generated based on the feature mapping results. Based on the fusion features and the fusion sub-features extracted from the task features and image features, fusion object features are generated, and pixel transformation is performed on the fusion object features to obtain a fusion virtual object.

[0095] In the specific execution process, firstly, feature extraction is performed on the task data to obtain task features, and image feature extraction is performed on multiple virtual objects to obtain multiple image features. Then, feature mapping is performed between the task features and the image features of multiple virtual objects. Based on the feature mapping results, object fusion rules for fusing multiple virtual objects are generated, and fusion features are generated based on the object fusion rules. Further, sub-features are extracted from text features and multiple image features according to the object fusion rules, and the extracted sub-features are concatenated to form fusion object features. Alternatively, the extracted sub-features are concatenated to form initial fusion features, and feature supplementation and feature reconstruction are performed on the initial fusion features to obtain fusion object features. Finally, pixel conversion is performed on the fusion object features to obtain the fusion object image as the fusion virtual object.

[0096] In addition, during the process of generating fusion features based on feature mapping results, the object fusion rules corresponding to the fusion features may also include the main virtual object and subordinate virtual objects among multiple virtual objects; Specifically, the primary and subordinate virtual objects among multiple virtual objects are determined in the following ways: the resource amount relationship between the resource amounts corresponding to the multiple virtual objects is determined, and the primary and subordinate virtual objects are determined based on the resource amount relationship and the resource characteristics contained in the theme role characteristics; or, the primary and subordinate virtual objects are determined based on the resource amounts corresponding to each of the multiple virtual objects.

[0097] The resource amount refers to the amount of resources that users need to pay to interact with the streamer, the live stream provider, or the live stream service provider by selecting virtual objects. The resources can be virtual resources or funds, such as carbon savings.

[0098] A primary virtual object refers to a virtual object that serves as the main reference or plays a major role when a virtual object is generated by merging multiple virtual objects; a subordinate virtual object refers to a virtual object that serves as a secondary reference or plays a minor role when a virtual object is generated by merging multiple virtual objects.

[0099] During the process of generating fused virtual objects, decisions can also be made regarding the resource amount of the fused virtual objects. For example, after obtaining fused virtual objects by pixel conversion of the features of the fused objects, the resource amount corresponding to the fused virtual objects can be determined based on the resource amount of each of the multiple virtual objects and the theme role field.

[0100] Specifically, in one optional implementation of this embodiment, the generation of fused virtual objects further includes: Extract the object resource features corresponding to the resource amounts of multiple virtual objects, and extract the resource features of the theme role field to obtain the theme resource features; Perform feature matching between object resource features and topic resource features, and determine resource weights based on the feature matching results; The resource amount corresponding to the merged virtual object is determined by the resource weight and the resource amount corresponding to each of the multiple virtual objects.

[0101] In the process of determining the resource amount of the merged virtual object based on the resource amount of each of the multiple virtual objects and the theme role field, the object resource features of the resource amount of each of the multiple virtual objects are first extracted, and the theme resource features are obtained by extracting the resource features of the theme role field. Then, the object resource features and theme resource features are matched, and the resource weight is determined based on the feature matching results. Further, the resource amount is calculated based on the resource weight and the resource amount of each of the multiple virtual objects to obtain the initial resource amount, and the initial resource amount is adjusted based on the theme resource features to obtain the resource amount of the merged virtual object.

[0102] In addition, in determining the resource amount corresponding to the merged virtual object, the resource amount can also be determined based on the resource amount of each virtual object and its relationship characteristics with the theme, master-slave relationship, and / or user characteristics. Specifically, determining the resource amount based on the resource amount of each virtual object and its relationship characteristics with the theme, master-slave relationship, and / or user characteristics is similar to the implementation method described above. User characteristics refer to the features obtained by feature extraction from user feature data.

[0103] In practical applications, during the generation of fused virtual objects, the generative model can further generate fused virtual objects that better match the relationships between the multiple virtual objects by using the master and subordinate virtual objects among the multiple virtual objects. Specifically, in one optional implementation provided in this embodiment, the generation of fused virtual objects by the generative model includes: The main virtual object and subordinate virtual objects are determined based on the image semantic features of multiple virtual objects, or based on the object type and / or user characteristics of multiple virtual objects. The main object features of the main virtual object and the subordinate object features of the subordinate virtual objects are fused to obtain fused features. Based on the task features of the task data, the fused features are adjusted and pixel transformed to obtain the fused virtual object.

[0104] Furthermore, during the generation of fused virtual objects, the generative model can also generate fused virtual objects that better match the relationships between the multiple virtual objects by using the master virtual object and subordinate virtual objects among the multiple virtual objects. Specifically, in another optional implementation provided in this embodiment, the generation of fused virtual objects by the generative model includes: User features and holiday features are mapped to image features of multiple virtual objects, and fused features are generated based on the feature mapping results. The fusion object features are generated based on the fusion features and the fusion sub-features extracted from the festival features and image features. Then, pixel transformation is performed on the fusion object features to obtain the fused virtual object. Here, the festival features refer to the features obtained by extracting features from the festival information of the current festival.

[0105] It should be noted that the various implementation methods provided above for constructing fused data based on multiple virtual objects and user feature data, and inputting the fused data into a generative model to generate fused virtual objects, can be combined in any form to form new implementation methods according to actual execution needs. Alternatively, one or more steps from different implementation methods can be combined to form new implementation methods suitable for actual execution scenarios. For example, generating fused virtual objects includes: determining the main virtual object and subordinate virtual objects based on the image semantic features, object types, and / or user features of multiple virtual objects; performing feature fusion on the main object features of the main virtual object and the subordinate object features of the subordinate virtual objects to obtain fused features, and performing feature adjustment and pixel transformation on the fused features based on the task features of the task data to obtain fused virtual objects; performing feature matching between the object resource features corresponding to the resource amounts of each of the multiple virtual objects and the theme resource features of the theme role field (or the task features of the task data), and determining the resource weights based on the feature matching results; and determining the resource amounts corresponding to the fused virtual objects based on the resource weights and the resource amounts corresponding to each of the multiple virtual objects.

[0106] In specific implementation, based on the above-mentioned server generating fused virtual objects by inputting fused data constructed based on multiple virtual objects and user feature data into a generative model, the server further returns the obtained fused virtual objects to the user terminal. Accordingly, here, the fused virtual objects returned by the server are received.

[0107] Step S706: Perform the display processing of the fused virtual object to conduct live interactive processing of the fused virtual object in cooperation with the server.

[0108] Based on the fused virtual object returned by the server, the fused virtual object is displayed here to enable live interactive processing of the fused virtual object in cooperation with the server. Specifically, the live interactive processing of the fused virtual object can involve the user confirming the interaction with the fused virtual object, and the user terminal sending an interaction confirmation instruction to the server based on the interaction confirmation. After receiving the interaction confirmation instruction, the server deducts resources from the user's resource account to perform the interactive processing of the fused virtual object.

[0109] For example, Figure 3 The interactive list shown is the live page of the live room accessed by the user through the user terminal. The interactive list consists of multiple virtual objects. The user selects virtual object 301 and virtual object 302 in the interactive list and initiates the fusion of virtual objects by triggering the "Merge" button configured on the live page. After the user terminal submits the virtual object 301 and virtual object 302 selected by the user in the interactive list to the server, the server queries the object fusion record of the virtual object fusion performed by the user. If the query result is empty, the user feature data is obtained. Based on virtual object 301, virtual object 302, and user feature data, fused data is constructed. This fused data is then input into a generative model to generate fused virtual objects. The generated fused virtual objects are returned to the user terminal, and the fused virtual objects displayed on the user terminal are as follows: Figure 4 As shown.

[0110] Furthermore, after returning the obtained fused virtual object to the user terminal, the live interaction processing of the user regarding the fused virtual object can also involve updating the fused virtual object. Specifically, in one optional implementation of this embodiment, the live interaction processing of the fused virtual object includes: If an object update request is received from a user terminal, the object update data corresponding to the object update request is generated and input into the generative model to adjust the fused virtual object, thereby obtaining the updated virtual object; Alternatively, if an object update request is received from a user terminal, update fusion data is constructed based on the object update request, multiple virtual objects, and user feature data. The update fusion data is then input into a generative model to generate fused virtual objects, thereby obtaining updated virtual objects.

[0111] Specifically, in one optional implementation of this embodiment, adjustments are made to the merged virtual object, including: Adjust the master-slave relationship between the master virtual object and the subordinate virtual object among the multiple virtual objects stored in the generative model, and extract the adjusted master object features and subordinate object features; Virtual objects are generated based on the characteristics of the subject role, user, main object, and subordinate object, and updated virtual objects are obtained.

[0112] Following the previous example, the user terminal is displayed. Figure 4 Based on the merged virtual object shown, if a user updates the merged virtual object by clicking the "Change" button, the updated virtual object returned by the server after the update is as follows: Figure 5 As shown; where, Figure 4 It is a fused virtual object generated by fusing virtual object 301 as the main virtual object and virtual object 302 as the subordinate virtual object; Figure 5 It is a fused virtual object generated by the server after adjusting the master-slave relationship between virtual objects 301 and virtual objects 302, with virtual object 302 as the master virtual object and virtual object 301 as the subordinate virtual object.

[0113] It should be noted that the process of the server constructing updated fusion data based on object update requests, multiple virtual objects, and user feature data, and then inputting the fusion data into the generative model to generate fused virtual objects, can be implemented in a similar way to the above process of constructing fusion data based on multiple virtual objects and user feature data and then inputting the fusion data into the generative model to generate fused virtual objects. For example, adjusting the fused virtual object includes: performing feature mapping between the update features of the object update request and the fused object features of the fused virtual object, generating update features based on the feature mapping results; generating update object features based on the update features and the update sub-features extracted from the fused object features, and performing pixel transformation on the update object features to obtain the updated virtual object; For example, adjusting the fused virtual object includes: performing feature mapping between user features and festival features and image features of multiple virtual objects, generating updated features based on the feature mapping results; generating updated object features based on the updated features and the fused sub-features extracted from the festival features and image features, and performing pixel transformation on the updated object features to obtain the updated virtual object; Alternatively, feature mapping can be performed between the festival features and the fusion object features of the fused virtual object, and updated features can be generated based on the feature mapping results; updated object features can be generated based on the updated features and the updated sub-features extracted from the festival features and / or fused object features, and the updated object features can be pixel-transformed to obtain the updated virtual object.

[0114] In this embodiment, when the server receives multiple virtual objects selected by the user terminal in the interaction list of the live broadcast room, it can perform anomaly verification on the live broadcast room. If the verification passes, it creates asynchronous interaction tasks corresponding to multiple virtual objects and adds the asynchronous interaction tasks to the asynchronous task queue. In this way, the asynchronous task queue can realize parallel virtual object interaction, that is, it can call the generative model in parallel to generate fused virtual objects, which helps to improve the response efficiency and real-time interaction of virtual object interaction in the live broadcast room.

[0115] Subsequently, during the server's processing of asynchronous interactive tasks in the asynchronous task queue, based on the generation of the fusion object or the updating of the virtual object, it queries whether the interaction status of the fusion object or the updating of the virtual object is normal. If so, it reads the resource amount of the fusion object or the updating of the virtual object and queries whether the resource amount in the accessing user's resource account is greater than or equal to the resource amount of the fusion object or the updating of the virtual object. If it is greater, it creates an object interaction record for the accessing user and configures the order status of the object interaction order in the object interaction record to the initialization state. Further, it calls the resource system corresponding to the accessing user's resource account to deduct resources. If the deduction is successful, it updates the order status of the object interaction order in the object interaction record to the paid state. If the deduction fails, it cancels the object interaction order and updates the order status of the object interaction order to the canceled state.

[0116] In summary, the live streaming processing method provided in this embodiment obtains multiple virtual objects selected by the user in the live streaming room during the process of the user accessing the live streaming room through the user terminal and submits them to the server. The server queries the object fusion record of the virtual objects fusion performed by the user. Based on the fused object record, fusion data can be constructed based on multiple virtual objects and the object fusion record. The fusion data is then input into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects that conform to the user's object fusion habits and preferences. Subsequently, the server can also return the fused virtual objects to the user terminal. Accordingly, the user terminal receives the fused virtual objects returned by the server and performs fused virtual object display processing to perform live interactive processing of fused virtual objects, thereby improving the interactive diversity of virtual objects interacting with the user in the live streaming room. If the server does not find an object fusion record, the server constructs fusion data based on multiple virtual objects and user feature data, and inputs the fusion data into a generative model to generate fusion virtual objects. The server then returns the fusion virtual objects to the user terminal. Correspondingly, the server receives the fusion virtual objects returned by the server and performs display processing on the fusion virtual objects to carry out live interactive processing of the fusion virtual objects, thereby enhancing the interactive fun of users interacting with virtual objects in the live broadcast room.

[0117] The following example uses a live streaming processing method provided in this embodiment in a live streaming platform scenario as an example, combined with... Figure 6 The network live streaming processing method provided in this embodiment will be further explained below. See [link to documentation]. Figure 6 The network live streaming processing method applied to live streaming platform scenarios includes the following steps.

[0118] Step S602: Obtain the multiple virtual objects selected by the user in the live broadcast room and submit them to the server.

[0119] Step S616: Receive the fused virtual object returned by the server.

[0120] Optionally, the fused virtual object is obtained by inputting fused data, which is constructed based on multiple virtual objects and user feature data, into a generative model to generate the fused virtual object.

[0121] Step S618: Perform the display processing of the merged virtual objects.

[0122] Step S620: Obtain the update operation of the accessing user for the merged virtual object, and submit the object update request to the server.

[0123] Step S628: Receive the updated virtual object returned by the server and perform the display processing of the updated virtual object.

[0124] It should be noted that any one or more of steps S602, S616 to S620, and S628 can be combined with any one or more of steps S702 to S706 to form a new implementation method according to the needs of implementation and deployment. In addition, any one or more technical features in steps S602, S616 to S620, and S628 can be selected and combined with any one or more technical features provided in steps S702 to S706 to form a new implementation method according to the actual deployment needs. Alternatively, any one or more technical features in steps S602, S616 to S620, and S628 can also be replaced with any one or more technical features provided in steps S702 to S706 to form a new implementation method according to the actual deployment needs. These will not be elaborated on here.

[0125] This specification provides an embodiment of a network live streaming processing device as follows: In the above embodiments, a method for processing live streaming is provided, and correspondingly, a device for processing live streaming is also provided, which will be described below with reference to the accompanying drawings.

[0126] Reference Figure 8This illustration shows a schematic diagram of an embodiment of a network live streaming processing device provided in this embodiment.

[0127] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.

[0128] This embodiment provides a network live streaming processing device, the device comprising: The object receiving module 802 is configured to receive multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room; The record query module 804 is configured to query the object fusion records of the virtual object fusion performed by the accessing user. If the query result is empty, the user feature data is obtained. The object generation module 806 is configured to construct fused data based on the multiple virtual objects and the user feature data, and input the fused data into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects; The object return module 808 is configured to return the fused virtual object to the user terminal for live interactive processing of the fused virtual object.

[0129] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0130] Another embodiment of the network live streaming processing device provided in this specification is as follows: In the above embodiments, another method for processing live streaming is provided, and correspondingly, another apparatus for processing live streaming is also provided, which will be described below with reference to the accompanying drawings.

[0131] Reference Figure 9 This illustration shows a schematic diagram of an embodiment of a network live streaming processing device provided in this embodiment.

[0132] Since the apparatus embodiments correspond to the method embodiments, the descriptions are relatively simple. For relevant parts, please refer to the corresponding descriptions of the method embodiments provided above. The apparatus embodiments described below are merely illustrative.

[0133] This embodiment provides a network live streaming processing device, the device comprising: The object submission module 902 is configured to retrieve multiple virtual objects selected by the user in the live broadcast room and submit them to the server. The object receiving module 904 is configured to receive the fused virtual object returned by the server; the fused virtual object is obtained by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model to generate the fused virtual object; The object display module 906 is configured to perform the display processing of the fused virtual object, so as to perform live interactive processing of the fused virtual object in cooperation with the server.

[0134] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0135] This specification provides the following server implementation example: Corresponding to the network live streaming processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a server for executing the network live streaming processing method provided above. Figure 10 This is a schematic diagram of the structure of a server provided for one or more embodiments of this specification.

[0136] This embodiment provides a server, including: like Figure 10As shown, device 1000 mainly consists of a communication interface 1002, a user interface 1004, a processor 1006, and a data storage 1008. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1010. The communication interface 1002 enables device 1000 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1002 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1002 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1002 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1002 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1004 includes receiving user input and providing output to the user. Therefore, the user interface 1004 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. The user interface 1004 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, the user interface 1004 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, the device 1000 may support remote access from other devices via communication interface 1002 or another physical interface (not shown). The user interface 1004 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. The user interface 1004 may also be configured as a display device for rendering or displaying text fragments.

[0137] Processor 1006 may include one or more general-purpose processors and / or dedicated processors. Data storage 1008 may include one or more volatile and / or non-volatile storage components, and may be integrated wholly or partially with processor 1006. Data storage 1008 may include removable and non-removable components.

[0138] Processor 1006 is capable of executing program instructions 1018 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 1008 to perform the various functions described herein. Data storage 1008 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1000, enable device 1000 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1018 by processor 1006 may result in processor 1006 using data 1012. For example, program instructions 1018 may include an operating system 1022 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1000 and one or more application programs 1020 (e.g., a browser, social application, or game application). Similarly, data 1012 may include operating system data 1016 and application data 1014. Operating system data 1016 is primarily accessible to operating system 1022, while application data 1014 is primarily accessible to one or more application programs 1020. Application data 1014 may reside in a file system visible or hidden by the user of device 1000. Application 1020 may communicate with operating system 1022 via one or more application programming interfaces (APIs). These APIs facilitate application 1020 reading and / or writing application data 1014, transmitting or receiving information via communication interface 1002, receiving or displaying information on user interface 1004, etc. In some terms, application 1020 may be simply referred to as an "app". Furthermore, application 1020 may be downloaded to device 1000 through one or more online app stores or app markets. However, applications may also be installed on device 1000 in other ways, such as through a web browser or a physical interface on device 1000 (e.g., a USB port).

[0139] In one specific embodiment, the server includes memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the server, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Receive multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room; Query the object fusion records of the user who performed virtual object fusion. If the query result is empty, retrieve the user's feature data. Based on the multiple virtual objects and the user feature data, fused data is constructed, and the fused data is input into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects; The merged virtual object is returned to the user terminal for live interactive processing of the merged virtual object.

[0140] This specification provides an example of a user terminal as follows: Corresponding to the other network live streaming processing method described above, based on the same technical concept, one or more embodiments of this specification also provide a user terminal for executing the other network live streaming processing method provided above. Figure 11 This is a schematic diagram of the structure of a user terminal provided for one or more embodiments of this specification.

[0141] This embodiment provides a user terminal, including: like Figure 11As shown, device 1100 mainly consists of a communication interface 1102, a user interface 1104, a processor 1106, and a data storage 1108. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 1110. The communication interface 1102 enables device 1100 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 1102 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 1102 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 1102 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 1102 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces. The user interface 1104 includes receiving user input and providing output to the user. Therefore, user interface 1104 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 1104 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 1104 may include software, circuitry, or other forms of logic capable of transmitting data to and receiving data from external user input / output devices. Additionally or alternatively, device 1100 may support remote access from other devices via communication interface 1102 or another physical interface (not shown). User interface 1104 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 1104 may also be configured as a display device for rendering or displaying text fragments.

[0142] Processor 1106 may include one or more general-purpose processors and / or dedicated processors. Data storage 1108 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 1106. Data storage 1108 may include removable and non-removable components.

[0143] Processor 1106 is capable of executing program instructions 1118 (e.g., compiled or uncompiled program logic and / or machine code) stored in data store 1108 to perform the various functions described herein. Data store 1108 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 1100, enable device 1100 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 1118 by processor 1106 may result in processor 1106 using data 1112. For example, program instructions 1118 may include an operating system 1122 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 1100 and one or more application programs 1120 (e.g., a browser, social application, or game application). Similarly, data 1112 may include operating system data 1116 and application data 1114. Operating system data 1116 is primarily accessible to operating system 1122, while application data 1114 is primarily accessible to one or more application programs 1120. Application data 1114 may reside in a file system visible or hidden to the user of device 1100. Application 1120 may communicate with operating system 1122 via one or more application programming interfaces (APIs). These APIs facilitate application 1120 reading and / or writing application data 1114, transmitting or receiving information via communication interface 1102, receiving or displaying information on user interface 1104, etc. In some terms, application 1120 may be simply referred to as an "app". Furthermore, application 1120 may be downloaded to device 1100 through one or more online app stores or app markets. However, applications may also be installed on device 1100 in other ways, such as through a web browser or a physical interface on device 1100 (e.g., a USB port).

[0144] In one specific embodiment, the server includes memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the server, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Retrieve multiple virtual objects selected by the user in the live stream and submit them to the server; The server returns a fused virtual object; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model. The fused virtual object is displayed and processed to enable live interactive processing of the fused virtual object in cooperation with the server.

[0145] This specification provides an embodiment of a computer-readable storage medium as follows: Corresponding to the above-described method for processing live streaming, and based on the same technical concept, one or more embodiments of this specification also provide a computer-readable storage medium.

[0146] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Receive multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room; Query the object fusion records of the user who performed virtual object fusion. If the query result is empty, retrieve the user's feature data. Based on the multiple virtual objects and the user feature data, fused data is constructed, and the fused data is input into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects; The merged virtual object is returned to the user terminal for live interactive processing of the merged virtual object.

[0147] It should be noted that the embodiments of a computer-readable storage medium described in this specification and the embodiments of a network live streaming processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0148] Another example of a computer program product provided in this specification is as follows: Corresponding to the other network live streaming processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another computer program product.

[0149] The computer-readable storage medium provided in this embodiment is used to store computer-executable instructions, which, when executed, implement the following process: Retrieve multiple virtual objects selected by the user in the live stream and submit them to the server; The server returns a fused virtual object; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model. The fused virtual object is displayed and processed to enable live interactive processing of the fused virtual object in cooperation with the server.

[0150] It should be noted that the embodiments of another computer program product described in this specification and the embodiments of another network live streaming processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0151] This specification provides an example of a computer program product as follows: Corresponding to the above-described method for processing live streaming, based on the same technical concept, one or more embodiments of this specification also provide a computer program product.

[0152] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Receive multiple virtual objects selected by the user terminal from the interaction list in the live broadcast room; Query the object fusion records of the user who performed virtual object fusion. If the query result is empty, retrieve the user's feature data. Based on the multiple virtual objects and the user feature data, fused data is constructed, and the fused data is input into a generative model to generate fused virtual objects, thereby obtaining fused virtual objects; The merged virtual object is returned to the user terminal for live interactive processing of the merged virtual object.

[0153] It should be noted that the embodiments of a computer program product described in this specification and the embodiments of a network live streaming processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0154] Another example of a computer program product provided in this specification is as follows: Corresponding to the other network live streaming processing method described above, based on the same technical concept, one or more embodiments of this specification also provide another computer program product.

[0155] A computer program product includes a computer program / instructions that, when executed by a processor, perform the following steps: Retrieve multiple virtual objects selected by the user in the live stream and submit them to the server; The server returns a fused virtual object; the fused virtual object is generated by inputting fused data constructed based on the multiple virtual objects and user feature data into a generative model. The fused virtual object is displayed and processed to enable live interactive processing of the fused virtual object in cooperation with the server.

[0156] It should be noted that the embodiments of another computer program product described in this specification and the embodiments of another network live streaming processing method described in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding method described above, and the repeated parts will not be described again.

[0157] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. For example, the device embodiment, equipment embodiment and computer-readable storage medium embodiment are all similar to the method embodiment, so the description is relatively simple. When reading the relevant content of the device embodiment, equipment embodiment and computer-readable storage medium embodiment, please refer to the description of the method embodiment.

[0158] While one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible execution order among many steps, and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims. This specification uses specific terms to describe embodiments of this specification. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0159] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0160] In the 1930s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement to the methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0161] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0162] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0163] For ease of description, the above apparatus is described by dividing it into various functional units. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in one or more software and / or hardware.

[0164] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0166] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0167] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0168] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0169] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0170] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0171] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising at least one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0172] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0173] The above description is merely an embodiment of this document and is not intended to limit the scope of this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.

Claims

1. A network live processing method, comprising: receiving a plurality of virtual objects selected by a user terminal in an interaction list of a live room; querying object fusion records of an access user for virtual object fusion, and obtaining user feature data if the query result is empty; constructing fusion data based on the plurality of virtual objects and the user feature data, and inputting the fusion data into a generative model to generate a fusion virtual object, to obtain the fusion virtual object; returning the fusion virtual object to the user terminal for live interaction processing of the fusion virtual object.

2. The network live processing method of claim 1, wherein the constructing fusion data based on the plurality of virtual objects and the user feature data comprises: writing the user feature data into a fusion task template to obtain task data, and associating the plurality of virtual objects with the task data to obtain the fusion data; wherein the fusion task template records at least one of the following: a theme role field for virtual object fusion, an input rule field, a generation task field for generating a fusion virtual object, and an output rule field.

3. The network live processing method of claim 2, wherein the generating a fusion virtual object comprises: performing feature mapping on task features of the task data and image features of the plurality of virtual objects, and generating fusion features according to the feature mapping result; generating fusion object features according to the fusion features and fusion sub-features extracted from the task features and the image features, and performing pixel conversion on the fusion object features to obtain the fusion virtual object.

4. The network live processing method of claim 3, wherein the generating a fusion virtual object further comprises: extracting object resource features corresponding to respective resource amounts of the plurality of virtual objects, and extracting theme resource features from the theme role field; performing feature matching on the object resource features and the theme resource features, and determining resource weights according to the feature matching result; determining resource amounts corresponding to the fusion virtual object according to the resource weights and the respective resource amounts of the plurality of virtual objects.

5. The network live processing method of claim 3, wherein the fusion rules corresponding to the fusion features include a master virtual object and a subordinate virtual object in the plurality of virtual objects; wherein the master virtual object and the subordinate virtual object in the plurality of virtual objects are determined in the following manner: determining resource amount relationships of resource amounts corresponding to the plurality of virtual objects, and determining the master virtual object and the subordinate virtual object according to the resource amount relationships and resource features included in the theme role features; or determining the master virtual object and the subordinate virtual object according to the respective resource amounts of the plurality of virtual objects.

6. The network live processing method of claim 1, wherein the fusion virtual object generation by the generative model comprises: determine the master virtual object and the subordinate virtual object according to image semantic features of the plurality of virtual objects, or determine the master virtual object and the subordinate virtual object according to object types of the plurality of virtual objects and / or user features; perform feature fusion on master object features of the master virtual object and subordinate object features of the subordinate virtual object to obtain fused features, and perform feature adjustment and pixel conversion on the fused features based on task features of the task data to obtain the fused virtual object.

7. The network live broadcast processing method according to claim 1, wherein the fused virtual object generated by the generative model comprises: performing feature mapping on user features and festival features and image features of the plurality of virtual objects, and generating fused features according to a feature mapping result; generating fused object features according to the fused features and fused sub-features extracted from the festival features and the image features, and performing pixel conversion on the fused object features to obtain the fused virtual object.

8. The network live broadcast processing method according to claim 1, wherein the live broadcast interaction processing of the fused virtual object comprises: if an object update request submitted by the user terminal is received, generating object update data corresponding to the object update request and inputting the object update data into the generative model to adjust the fused virtual object, and obtaining an updated virtual object.

9. The network live broadcast processing method according to claim 8, wherein the adjustment of the fused virtual object comprises: adjusting master-slave relationships of master virtual objects and subordinate virtual objects in the plurality of virtual objects stored in the generative model, and extracting adjusted master object features and subordinate object features; generating a virtual object based on theme role features, user features, and the master object features and the subordinate object features, and obtaining the updated virtual object.

10. The network live broadcast processing method according to claim 1, further comprising, after the object fusion record operation of querying a user to perform virtual object fusion is executed: if an object fusion record is queried, constructing fused data based on the plurality of virtual objects and the object fusion record, and inputting the fused data into the generative model to generate a fused virtual object, and obtaining the fused virtual object.

11. The network live broadcast processing method according to claim 10, wherein the fused virtual object generation comprises: performing feature mapping on text features of the object fusion record and image features of the plurality of virtual objects, and generating fused features according to a feature mapping result; generating fused object features according to the fused features and fused sub-features extracted from the text features and the image features, and performing pixel conversion on the fused object features to obtain a fused object image.

12. The network live broadcast processing method according to claim 1, further comprising, after the step of receiving a plurality of virtual objects selected by a user terminal in an interaction list of a live broadcast room is executed, and before the object fusion record operation of querying a user to perform virtual object fusion is executed: reading fused object identifiers corresponding to the plurality of virtual objects and returning the fused object identifiers to the user terminal. If the user terminal submits a fusion confirmation instruction, the object fusion record operation of the user terminal is performed; If the user terminal submits a switching instruction, the updated object identifier corresponding to the multiple virtual objects is generated and returned to the user terminal.

13. The network live processing method of claim 12, wherein the fusion object identifier is obtained by inputting the multiple virtual objects and the fusion task template into the generative model to generate a fusion virtual object identifier. The updated object identifier is obtained by adjusting the fusion object identifier, or by adjusting the object relationship of the multiple virtual objects and inputting the adjusted multiple virtual objects and the fusion task template into the generative model to generate a fusion virtual object identifier.

14. A network live processing method, comprising: Obtaining multiple virtual objects selected by an access user in a live room and submitting to a server; Receiving a fusion virtual object returned by the server; The fusion virtual object is obtained by inputting fusion data constructed based on the multiple virtual objects and user feature data into a generative model to generate a fusion virtual object; Performing display processing of the fusion virtual object to perform live interaction processing of the fusion virtual object with the server.

15. A network live processing apparatus, comprising: An object receiving module configured to receive multiple virtual objects selected by a user terminal in an interaction list in a live room; A record query module configured to query an object fusion record of a virtual object fusion by an access user, and obtain user feature data if the query result is empty; An object generation module configured to construct fusion data based on the multiple virtual objects and the user feature data, and input the fusion data into a generative model to generate a fusion virtual object, and obtain a fusion virtual object; An object returning module configured to return the fusion virtual object to the user terminal to perform live interaction processing of the fusion virtual object.

16. A network live processing apparatus, comprising: An object submitting module configured to obtain multiple virtual objects selected by an access user in a live room and submit to a server; An object receiving module configured to receive a fusion virtual object returned by the server; The fusion virtual object is obtained by inputting fusion data constructed based on the multiple virtual objects and user feature data into a generative model to generate a fusion virtual object; An object display module configured to perform display processing of the fusion virtual object to perform live interaction processing of the fusion virtual object with the server.

17. A server, comprising: A processor; And a memory configured to store computer executable instructions which, when executed, cause the processor to: Receive multiple virtual objects selected by a user terminal in an interaction list in a live room; Query an object fusion record of a virtual object fusion by an access user, and obtain user feature data if the query result is empty; construct fusion data based on the plurality of virtual objects and the user feature data, and input the fusion data into a generative model to generate a fusion virtual object; return the fusion virtual object to the user terminal for live interaction processing of the fusion virtual object. 18.A user terminal, comprising: a processor; and a memory configured to store computer executable instructions that, when executed, cause the processor to: obtain a plurality of virtual objects selected by a user in a live room and submit to a server; receive a fusion virtual object returned by the server; the fusion virtual object being obtained by inputting fusion data constructed based on the plurality of virtual objects and user feature data into a generative model to generate a fusion virtual object; perform display processing of the fusion virtual object to cooperate with the server to perform live interaction processing of the fusion virtual object. 19.A computer readable storage medium for storing computer executable instructions that, when executed, implement the steps of the method of claim 1 or claim 14.