Data processing method and device for virtual scene
By using multimodal data fusion and historical data modeling, the parameters of the virtual scene are dynamically adjusted, which solves the problem of poor adaptability of the virtual scene and improves the user experience and immersion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING QIMIAO KINGDOM TECHNOLOGY CO LTD
- Filing Date
- 2025-09-25
- Publication Date
- 2026-06-02
Smart Images

Figure CN121371614B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data processing method and apparatus for virtual scenes. Background Technology
[0002] With the development of computer technology, electronic devices can realize richer and more vivid virtual scenes. A virtual scene refers to a digital scene outlined by a computer through digital communication technology. Users can obtain a fully virtualized experience (such as virtual reality) or a partially virtualized experience (such as augmented reality) in the virtual scene, and can interact with various objects in the virtual scene or control the interaction between various objects in the virtual scene to obtain feedback.
[0003] In related technologies, virtual scenes often adopt static design, that is, using fixed scene parameters (such as game difficulty, game music rhythm, etc.). However, this makes it difficult to adapt to individual user differences, affecting the user's experience in the virtual scene. Summary of the Invention
[0004] This application provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for virtual scenes, which can improve the adaptability of virtual scenes to individual users and enhance user experience and immersion.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides a data processing method for a virtual scene, including:
[0007] Obtain object data of multiple modalities of the target object participating in the virtual scene, and extract object features of the object data of each modality;
[0008] The object features of multiple modalities are fused to obtain fused object features;
[0009] Obtain historical data of the target object's participation in the virtual scene, and construct an object profile of the target object based on the characteristics of the fused object and the historical data;
[0010] The environmental data of the target object is obtained, and the target scene in which the target object is located is predicted based on the environmental data.
[0011] By combining the object profile and the target scene, the scene parameters of the virtual scene are adjusted.
[0012] This application embodiment also provides a data processing device for a virtual scene, including:
[0013] An extraction module is used to acquire object data of multiple modalities of target objects participating in a virtual scene, and to extract object features of the object data of each modality;
[0014] A fusion module is used to fuse the object features of multiple modalities to obtain fused object features;
[0015] A construction module is used to acquire historical data of the target object's participation in the virtual scene, and to construct an object profile of the target object based on the characteristics of the fused object and the historical data;
[0016] The prediction module is used to acquire environmental data of the target object and, based on the environmental data, predict the target scene in which the target object is located.
[0017] The adjustment module is used to adjust the scene parameters of the virtual scene by combining the object profile and the target scene.
[0018] This application also provides an electronic device, including:
[0019] Memory is used to store executable instructions for a computer;
[0020] The processor, when executing computer-executable instructions stored in the memory, implements the data processing method for the virtual scene provided in the embodiments of this application.
[0021] This application also provides a computer-readable storage medium storing computer-executable instructions or computer programs, which, when executed by a processor, implement the data processing method for a virtual scene provided in this application.
[0022] This application also provides a computer program product, including computer-executable instructions or a computer program, which, when executed by a processor, implements the data processing method for a virtual scene provided in this application.
[0023] The embodiments of this application have the following beneficial effects:
[0024] Applying the embodiments described above, firstly, object data of the target object participating in the virtual scene in multiple modalities is acquired, and object features of the object data in each modality are extracted. Then, the object features of multiple modalities are fused to obtain fused object features. Secondly, historical data of the target object's participation in the virtual scene is acquired, and an object profile of the target object is constructed based on the fused object features and historical data. Thirdly, environmental data of the target object is acquired, and the target scene in which the target object is located is predicted based on the environmental data. Finally, the scene parameters of the virtual scene are adjusted by combining the object profile and the target scene. Thus, by using multimodal data fusion and historical data modeling to accurately construct an object profile, and by predicting the target scene in which the object is located using environmental data, the virtual scene parameters can be dynamically adjusted based on the object profile and the target scene. This significantly improves the adaptability of the virtual scene to the individual user, enhancing the user experience and immersion. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the architecture of the data processing system for virtual scenes provided in the embodiments of this application;
[0026] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0027] Figure 3A This is a first flowchart illustrating the data processing method for a virtual scene provided in an embodiment of this application;
[0028] Figure 3B This is a schematic diagram of the second process of the data processing method for virtual scenes provided in the embodiments of this application;
[0029] Figure 3C This is a schematic diagram of the third process of the data processing method for a virtual scene provided in the embodiments of this application;
[0030] Figure 3D This is a schematic diagram of the fourth process of the data processing method for a virtual scene provided in the embodiments of this application;
[0031] Figure 3E This is a schematic diagram of the fifth process of the data processing method for a virtual scene provided in the embodiments of this application;
[0032] Figure 4 This is a schematic diagram of the object state prediction process provided in an embodiment of this application;
[0033] Figure 5 This is a schematic diagram of the process for determining parameter adjustment values provided in the embodiments of this application.
[0034] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0037] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0038] In this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of a larger module or unit that includes the functionality of the module or unit.
[0039] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0040] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0041] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0042] 1) Client: An application running on an electronic device that provides various services, such as a client that supports virtual scenes (like game scenes).
[0043] 2) In response to, used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed may be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0044] 3) A virtual scene is a virtual scene displayed (or provided) by an application while it is running on a terminal. This virtual scene can be a simulation of the real world, a semi-simulated / semi-fictional virtual environment, or a purely fictional virtual environment. A virtual scene can be any of the following: two-dimensional, 2.5-dimensional, or three-dimensional. For example, a virtual scene can include the sky, land, and ocean, and the land can include environmental elements such as deserts, cities, and houses. The virtual scene can be displayed from a first-person perspective (e.g., the user plays as a virtual object in the game); a third-person perspective (e.g., the user chases after a virtual object in the game); or a bird's-eye view. These perspectives can be switched arbitrarily.
[0045] 4) Scene parameters are the specific parameters of the virtual scene, used to depict the virtual scene experience, gameplay, performance, etc., and can be dynamically adjusted to adapt to the user's state. For example, scene parameters include, but are not limited to: player character parameters (such as the density and number of enemy players, the attack power, movement speed, health, skill strength, etc. of friendly players), non-player character parameters (such as the behavior strategies of NPCs and AI characters, such as attack frequency, skill release probability, attack range, etc.), interface parameters of the human-computer interaction interface of the virtual scene (such as the number and size of interface buttons, the complexity of UI operations, interface transparency, information display density, etc.), audio parameters of the virtual scene (such as music rhythm, ambient sound volume, sound effect tension), visual parameters of the virtual scene (such as visual color, brightness, style, complexity of visual effects), scene guidance parameters of the virtual scene (such as the display frequency, display position, prominence, and display timing of guidance and prompt information), interaction parameters of the virtual scene (such as game difficulty, interaction cooldown time, interaction success rate, skill cooldown time, game resource refresh frequency, game resource acquisition amount, game scene complexity, game scene interference factors (such as bad weather, terrain traps, etc.), etc.), and device performance adaptation parameters (such as graphics quality (texture resolution, model accuracy), frame rate limit).
[0046] This application provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product for virtual scenes, which can improve the adaptability of virtual scenes to individual users and enhance user experience and immersion. The following is a detailed description of the embodiments of this application based on the above explanation of the terms and concepts used.
[0047] The following describes the data processing system for virtual scenes provided in the embodiments of this application. See also: Figure 1 , Figure 1 This is a schematic diagram of the architecture of a virtual scene data processing system provided in an embodiment of this application. To support an exemplary application, the virtual scene data processing system 100 includes: a server 200, a network 300, and a terminal 400. The terminal 400 is connected to the server 200 via the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both, using wireless or wired links for data transmission.
[0048] Here, terminal 400 (e.g., running a client that supports virtual scenes (such as game scenes)) responds to a join command for a virtual scene triggered by a target object (such as a game player) and sends a request to server 200 to obtain scene data of the virtual scene; server 200 receives the request from terminal 400; in response to the request, it returns the scene data of the virtual scene to terminal 400; terminal 400 receives the scene data of the virtual scene returned by server 200; it renders the scene data to obtain the virtual scene and displays it. In this way, the target object can participate in the virtual scene and control the virtual object of the target object to interact in the virtual scene (such as fighting against the virtual objects of other game players); at the same time, the server 200 obtains object data of the target object participating in the virtual scene in multiple modalities and extracts the object features of the object data of each modality; the object features of multiple modalities are fused to obtain fused object features; the server 200 obtains the historical data of the target object participating in the virtual scene and constructs the object profile of the target object based on the fused object features and historical data; the server 200 obtains the environmental data of the target object and predicts the target scene in which the target object is located based on the environmental data; the server 200 combines the object profile and the target scene to determine the adjustment strategy of the scene parameters of the virtual scene (such as reducing the game difficulty by 20%, reducing the density of non-player characters by 10%, adjusting the rhythm of the game sound effects to a soothing rhythm, etc.) and returns the adjustment strategy to the terminal 400; the terminal 400 receives the adjustment strategy and adjusts the scene parameters of the virtual scene based on the adjustment strategy.
[0049] The virtual scene data processing method provided in this application embodiment is implemented by an electronic device. For example, it can be implemented by a terminal alone, by a server alone, or by a terminal and a server working together. The electronic device implementing the virtual scene data processing method provided in this application embodiment can be various types of terminals or servers. The server (e.g., server 200) can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal (e.g., terminal 400) can be a laptop, tablet, desktop computer, smartphone, smart voice interaction device (e.g., smart speaker), smart home appliance (e.g., smart TV), smartwatch, vehicle terminal, wearable device, virtual reality (VR) device, aircraft, etc., but is not limited to these. The terminal and server can be connected directly or indirectly through wired or wireless communication, and this application embodiment does not impose any restrictions on this.
[0050] In some embodiments, the terminal or server can implement the data processing method for the virtual scene provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as game APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.
[0051] The following describes an electronic device for implementing a data processing method for virtual scenes, as provided in an embodiment of this application. See also... Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 500 provided in this embodiment can be a terminal or a server. Figure 2As shown, electronic device 500 includes at least one processor 510, memory 550, at least one network interface 520, and user interface 530. The various components in electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2 The general labeled all buses as Bus System 540.
[0052] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0053] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0054] Memory 550 may be removable, non-removable, or a combination thereof. Memory 550 may include one or more storage devices physically located away from processor 510. Memory 550 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.
[0055] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0056] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0057] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0058] Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530;
[0059] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.
[0060] In some embodiments, the data processing device for virtual scenes provided in this application can be implemented in software. Figure 2 A data processing device 555 for a virtual scene stored in memory 550 is shown. It can be software in the form of programs and plug-ins, including the following software modules: extraction module 5551, fusion module 5552, construction module 5553, prediction module 5554, and adjustment module 5555. These modules are logical and can therefore be arbitrarily combined or further divided according to the functions they implement. The functions of each module will be described below.
[0061] The following describes the data processing method for virtual scenes provided in the embodiments of this application. As mentioned above, the data processing method for virtual scenes provided in the embodiments of this application is implemented by an electronic device, such as a server or terminal alone, or a server and terminal working together. Therefore, the executing entity of each step will not be described again below. See Figure 3A , Figure 3A This is a first flowchart illustrating the data processing method for a virtual scene provided in this application embodiment. The data processing method for a virtual scene provided in this application embodiment includes:
[0062] Step 101: Obtain object data for multiple modalities of the target object participating in the virtual scene, and extract object features from the object data of each modality.
[0063] For step 101, for example, the virtual scene can be a game scene, and the target object can be a user participating in the virtual scene (such as a player participating in a game). Here, the object data of multiple modalities of the target object is first obtained. For example, these multiple modalities include, but are not limited to, visual modalities, speech modalities, physiological modalities, and environmental modalities. The object data in the visual modal mainly consists of visual information acquired through devices such as cameras, including but not limited to facial images, full-body images, hand images, and leg images of the target object. The object data in the speech modal mainly consists of sound signals collected through microphones, including but not limited to the speech data of the target object. The object data in the physiological modal mainly consists of human physiological indicator data acquired through various physiological sensors (such as heart rate monitoring devices, EEG acquisition devices, skin conductance bracelets, smartwatches, etc.), including but not limited to the target object's heart rate and skin conductivity. The object data in the environmental modal mainly consists of various data related to the environment in which the virtual scene runs, which can be collected through devices (such as mobile phones), ambient light sensors, and temperature sensors, including but not limited to lighting data (such as the light intensity around the device supporting the virtual scene), noise data (the noise level around the device supporting the virtual scene), and geographic location data (such as GPS data of the device supporting the virtual scene). In practical applications, the above environmental data can also include environmental context data of the virtual scene (such as the device performance, network status, and usage time of the device displaying the virtual scene). After acquiring object data from multiple modalities, the object features of each modality are extracted. In practical applications, object data may also include operational behavior data of the target object (such as operational behavior data collected via gyroscopes / accelerometers).
[0064] In some embodiments, step 101, "extracting object features from object data of each modality," can be achieved by performing the following steps: if the modality is a visual modality and the object data includes an object image sequence of the target object, extract local image features of each object image in the object image sequence, and extract image temporal features between the local image features of the object image sequence, using the image temporal features as object features; if the modality is a speech modality and the object data includes object speech of the target object, extract emotional frequency band features, prosodic features, and spectral features of the object speech, and use all of these as object features; if the modality is a physiological modality and the object data includes a physiological data sequence of the target object, extract local physiological features of the physiological data sequence in each time window, and extract physiological temporal features between the local physiological features of the physiological data sequence, using the physiological temporal features as object features; if the modality is an environmental modality and the object data includes an environmental data sequence of the target object, perform temporal embedding processing on the environmental data sequence to obtain environmental temporal features, and map the environmental temporal features to environmental label features, using the environmental label features as object features.
[0065] (1) If the modality is visual, obtain the object image sequence of the target object. This object image sequence is a time series, which includes multiple object images (such as facial images) arranged in chronological order. Extract the local image features of each object image in the object image sequence. For example, extract the local image features of each object image through a convolutional neural network (such as ResNet18). This allows for the gradual extraction of features through multiple layers of convolutional kernels in the convolutional neural network. That is, first capture pixel-level details (such as the edge of eyebrows, the outline of the corners of the mouth, and the opening and closing state of the eyes), and then integrate the pixel-level details to form local image features (such as expression-related local features: furrowed brows, upturned corners of the mouth, wide eyes, etc.). Extracting temporal features from multiple local features of an object's image sequence is crucial. For example, a Long Short-Term Memory (LSTM) network can be used to extract temporal features between multiple local features. If the object image is a facial image, these temporal features not only contain the facial expression details of a single image but also integrate temporal contextual information such as the trend, speed, and duration of expression changes. For instance, an expression changing from calm to agitated (transition time 2 seconds, agitation lasting 5 seconds) more accurately reflects the user's real-time emotional state than a single local feature. Thus, accurately extracting local facial expression features (eyebrow, eye, and mouth movements) from each frame of a facial image using a convolutional neural network forms the basis for subsequent analysis. By concatenating the details (i.e., local image features) of multiple frames of facial images using LSTM, the patterns of facial expression changes over time can be uncovered, capturing key dynamic information such as micro-expressions and emotional shifts to accurately express the user's emotions. Finally, these temporal features are used as object features.
[0066] (2) If the modality is a speech modality, obtain the speech of the target object and convert the speech of the target object into a time spectrum, for example, by converting the speech of the target object into a time spectrum through short time Fourier transform (STFT); extract the emotional frequency band features from the time spectrum, for example, by using a convolutional neural network (such as a one-dimensional convolutional neural network) or a recurrent neural network, so as to extract the emotional frequency band features from the time spectrum. The emotional frequency band features may include key features such as high frequency noise intensity, speech rate change rate, and pitch tremor amplitude.
[0067] Simultaneously, prosodic features can be extracted from the target speech, including fundamental frequency (F0), speech rate, pause duration, and sound intensity. Fundamental frequency is the basic frequency of vocal cord vibration, reflecting pitch (e.g., F0 is high when excited, and low when depressed). The core of extraction is analyzing the periodic vibration patterns of the target speech's time-domain signal, such as through autocorrelation or cepstral methods. Speech rate is the number of syllables per unit time, which can be achieved through endpoint detection of the time-domain signal. For example, first locate the effective speech segments (excluding silent segments), then count the number of syllables or words within the effective speech segments, and divide by the time length to obtain the speech rate. Pause duration is the silence segment in the target speech, which can be determined by the intensity threshold of the time-domain signal. For example, if the speech signal intensity is below the threshold and the duration exceeds 0.1 seconds, it is considered a pause, and the pause duration is then calculated. Sound intensity is the loudness of the sound, which can be determined by calculating the average of the squares of the amplitude of the time-domain signal, and then determining the sound intensity based on the average value (e.g., the larger the average value, the greater the sound intensity). Simultaneously, spectral features of the target speech can be extracted from its time-spectrum graph. These spectral features can be Mel-frequency cepstral coefficients (MFCCs). In this way, emotional frequency band features, prosodic features, and spectral features are all used as target features.
[0068] (3) If the modality is a physiological modality, obtain the physiological data sequence of the target object. This physiological data sequence is a time series, which includes multiple physiological data (such as heart rate, skin conductance, etc.) arranged in chronological order. First, the physiological data of each time window can be extracted from the physiological data sequence by sliding time windows (such as the window length of the sliding time window is 4 seconds). For example, the physiological data of time window 1-4 seconds; then, feature extraction is performed on the physiological data of each time window to obtain local physiological features. For example, for each time window, at least one preset statistical indicator (such as mean, coefficient of variation, peak change rate, etc.) can be used to statistically analyze the physiological data of the time window to obtain the statistical results of each statistical indicator. The statistical results of the at least one statistical indicator are used as the local physiological features of the time window. Here, the mean is the average of all physiological data within the time window, reflecting the baseline physiological level; the coefficient of variation = standard deviation / mean, where the standard deviation is the standard deviation of all physiological data within the time window, reflecting the stability of signal fluctuations; the peak rate of change = (peak value of physiological data within the time window - trough value of physiological data within the time window) / window length of the time window, reflecting the speed of physiological state changes. Furthermore, since the local physiological features of each time window are isolated and cannot reflect the changing trends between different time windows, physiological temporal features are extracted from multiple local physiological features of the physiological data sequence, thus using the extracted physiological temporal features as object features. For example, a bidirectional gated recurrent unit (Bi-GRU) encoder can be used to process multiple local physiological features, mining the changing patterns of multiple local physiological features in the time dimension to obtain physiological temporal features.
[0069] (4) If the modality is an environment modality, obtain the environmental data sequence of the target object. This environmental data sequence is a time series, which includes multiple environmental data arranged in chronological order (such as light data, noise data, geographical location data, etc.). First, normalize the environmental data in the environmental data sequence to obtain a normalized environmental data sequence (including the normalized environmental data obtained by normalizing each environmental data). Then, embed the normalized environmental data sequence to obtain environmental time series features. For example, perform time embedding on the normalized environmental data sequence to obtain environmental time series features carrying time information. Specifically, assign a time embedding vector to each time step (such as 1 second as a time step). The time embedding vector is generated by encoding the time position (such as the 1st second, the 2nd second, etc.) and the time interval (such as the time interval between the current time step and the previous time step). Concatenate the time embedding vector with the feature vector of the normalized environmental data of the time step to obtain the environmental time series features carrying time information. In this way, through time embedding, it is possible to distinguish between stable scenes and dynamic scenes and avoid misjudging environmental label features due to fluctuations in instantaneous environmental data (such as sudden light flicker). Continuing, the temporal features of the environment are mapped to environmental label features, which are labels for the environment, such as nighttime environment, outdoor environment, noisy environment, quiet environment, etc. Specifically, the temporal features of the environment can be mapped using neural networks (such as multilayer perceptrons) to obtain environmental label features, which have fewer dimensions than the temporal features. In this way, scene parameters (such as UI style (brightness, color scheme), music rhythm, volume, prompt frequency, difficulty, etc.) are adjusted based on environmental data, making the presentation and rhythm of the virtual scene more matched with the environment. This improves the immersive experience of the virtual scene matching the environment, ensuring that the virtual scene always performs at an optimal level of immersion in different environments, thus meeting the scene perception requirements of mobile virtual scenes.
[0070] Applying the above embodiments, 1) each modality extracts object features through its own object feature extraction network, which can accurately adapt to the data characteristics of each modality (such as the spatiality of images and the temporality of speech), retain core information to the greatest extent, and can specifically optimize the feature extraction effect, improve the feature quality of each modality, lay a solid foundation for subsequent cross-modal fusion, make it easier for features of different modalities to be aligned and associated in subsequent stages, and thus enhance the expressive power and accuracy of the overall features after multimodal fusion; and can flexibly adapt to the sensor acquisition capabilities of different devices (such as ordinary mobile phone camera → simplified CNN to extract facial expression features, smartwatch heart rate → simplified Bi-GRU to extract physiological features), avoid device performance limitations on state recognition capabilities, and break free from device shackles. 2) It extracts features from multiple modalities (visual modality, speech modality, physiological modality, environmental modality, etc.), which can integrate multi-dimensional information such as vision, hearing, physiological signals, and environmental context to make up for the information limitations of a single modality; it enhances the richness and robustness of features, and can supplement information through other modalities when single-modal data is missing or interfered with, thus ensuring the integrity of information; it provides a comprehensive basis for subsequent multimodal fusion and accurate recognition of object states (such as emotions, attention, etc.), making the dynamic adjustment of scene parameters in virtual scenes more in line with the user's real experience, and greatly improving user personalization and accuracy.
[0071] Step 102: Fuse the object features of multiple modalities to obtain fused object features.
[0072] In step 102, after obtaining object features from multiple modalities, these features are fused to obtain fused object features. This fusion of object features from multiple modalities fully leverages the advantages of each modality (such as visual images capturing facial expressions, voice conveying tone, physiological data reflecting internal states, and environmental data reflecting the user's location), complementing the shortcomings of a single modality. This results in a more comprehensive and accurate capture of object states, significantly improving the accuracy and robustness of object state recognition. Furthermore, it reduces reliance on a single device, providing a more reliable multi-dimensional basis for the dynamic adjustment of scene parameters in virtual scenes.
[0073] In some embodiments, see Figure 3BStep 102, "Fusing object features from multiple modalities to obtain fused object features," can be achieved by executing the following steps 1021-1023: Step 1021, for each first modality among the multiple modalities, using the object features of the first modality as the query vector and the object features of at least one second modality as the key vector and value vector, attention processing is performed to obtain the attention weight between each second modality and the first modality. The first modality is any modality among the multiple modalities, and the second modality is a modality among the multiple modalities that is different from the first modality; Step 1022, for each first modality, according to the attention weight of each second modality, the object features of at least one second modality are weighted and fused to obtain the first fused feature, and the first fused feature and the object features of the first modality are fused to obtain the second fused feature of the first modality; Step 1023, the second fused feature of each first modality is mapped to the shared semantic space to obtain the fused object features.
[0074] For step 1021, for each first modality (which can be any modality among the multiple modalities), the object features of the first modality (e.g., the visual modality) are used as the query vector, and the object features of each second modality (a modality different from the first modality, such as the speech modality) are used as the key and value vectors. An attention mechanism is used to calculate the association weight between the first modality and each second modality, i.e., the attention weight. A higher attention weight indicates a stronger semantic association between the two modalities. For example, the semantic association between the "frowning feature" in the visual modality and the "increased tone feature" in the speech modality is strong, resulting in a high attention weight. In practical applications, a cross-modal attention layer can be introduced to implement the above attention processing. In practical applications, a fusion bottleneck token can be introduced, which enables attention processing to be implemented around the fusion bottleneck token. That is, for the object features of each modality, the fusion bottleneck token is used to filter features, retaining the core semantics that are most relevant and critical to the object state, thereby achieving feature compression and discarding redundant information. In this way, the object features filtered by the fusion bottleneck token are used to participate in feature fusion, avoiding the influence of redundant information and improving the accuracy of subsequent object state recognition.
[0075] For step 1022, firstly, for each first modality, according to the attention weight of each second modality, the object features of all second modalities are weighted and summed to obtain the first fused feature, i.e., the first fused feature = attention weight 1 * object feature of second modality 1 + attention weight 2 * object feature of second modality 2 + ... + attention weight N * object feature of second modality N, where N is the number of second modalities. Then, the first fused feature and the object features of the first modalities are fused to obtain the second fused feature of the first modality. The fusion method can be concatenation or residual addition, which is not limited here. Thus, the first fused feature is a fusion of cross-modal information, and the second fused feature is based on the fusion of cross-modal information and incorporates its own modal information, so that the second fused feature retains the core information of its own modality and incorporates the key information of related modalities.
[0076] For step 1023, the second fusion feature of each first modality is mapped to the shared semantic space to obtain the fusion object feature. For example, the second fusion feature of each first modality can be mapped to the shared semantic space through linear transformation or neural network. The second fusion features of different modalities may have different dimensions (e.g., 384 dimensions for images, 256 dimensions for speech), different time scales, and inconsistent semantic scales (images focus on visual details, speech focuses on temporal changes). The shared semantic space is a neutral feature space (e.g., unified to 128 dimensions), which makes the features of different modalities comparable and semantically consistent in the shared semantic space. For example, a state of tension corresponds to a specific region in the shared semantic space, regardless of whether the features come from the visual modality or the speech modality. Multiple second fusion features mapped to the shared semantic space can be averaged or weighted to obtain the fusion object feature. In this way, the fusion object feature contains the key information of all modalities, has a clear semantic meaning in the shared space, and eliminates modal differences, and can be directly used for subsequent processing.
[0077] Applying steps 1021-1023 above, 1) through cross-modal attention mechanisms and shared semantic space mapping, the semantic association strength between different modalities (such as visual modalities, speech modalities, physiological modalities, etc.) can be accurately quantified, allowing information from strongly associated modalities to fully interact; weighted fusion of associated modal features and their own features enables each modality to retain core information while incorporating information from complementary modalities, achieving cross-modal fusion, enhancing feature expressiveness, and possessing fault tolerance and dynamic adaptation capabilities; mapping multimodal features to the shared semantic space eliminates modal differences, obtaining unified and semantically aligned fused features, providing more comprehensive and accurate input for subsequent object state recognition (such as emotion, focus), improving state recognition accuracy and robustness, and helping to more accurately adapt the dynamic adjustment of scene parameters in virtual scenes to users. 2) It can perform cross-modal spatiotemporal feature extraction and fusion of image data, speech data, physiological data, and environmental data, far exceeding traditional splicing or weighting methods, significantly improving the accuracy and robustness of object state recognition, and effectively supporting the recognition of object states in complex environments.
[0078] See Figure 4 The object state prediction process includes: 1. Collecting object data from multiple modalities, including: image data (image sequences), speech data (audio frames), physiological data (heart rate, skin conductance signals), and environmental data. 2. Extracting object features from each modality, including: extracting features from image data using CNN+LSTM; extracting features from speech data using RNN or 1D CNN; extracting features from physiological data using STFT+Bi-GRU; and obtaining features from environmental data using MLP temporal embedding encoding. 3. Fusing the features from multiple modalities using a Cross-Modal Attention Fusion Layer to obtain fused object features. 4. Based on the fused object features, predicting the object state using a temporal modeling and prediction decoder, such as player state (emotion, focus level, fatigue warning, cognitive load level, etc.).
[0079] Step 103: Obtain historical data of the target object's participation in the virtual scene, and construct an object profile of the target object based on the fused object characteristics and historical data.
[0080] For step 103, firstly, historical data of the target object's participation in the virtual scene is obtained. For example, a cloud database is established to store player's completion records, number of failures, emotional curves, and other behavioral trajectory data. Specifically, the behavioral trajectory data includes: (1) game interaction data: completion records (such as completion time, level selection), number of failures (such as the failure frequency of a certain type of level), operation behavior (such as daily operation volume, key press frequency, dwell time), etc.; (2) state time series data: emotional curves (such as emotional changes within each game duration, from calm → excitement → irritability), physiological state time series (such as heart rate change curve), etc. In this way, the lack of personalization caused by relying solely on the immediate state (such as the current object state (such as emotion)) is avoided. By mining the stable characteristics of players through long-term behavioral trajectories (such as a long-term preference for exploration rather than combat, and irritability after 3 failures), a historical basis is provided for constructing a dynamic object profile of the target object. Then, based on the fusion of object characteristics and historical data, an object profile of the target object is constructed. Among them, the fused object features represent the current real-time state of the target object. By fusing the current real-time state and historical data, a dynamic profile of the target object is constructed, so that the object profile includes both the current real-time state and historical data, thereby improving the accuracy of dynamic adjustment of scene parameters based on the object profile. This makes the dynamic adjustment of scene parameters more in line with the user's real experience and greatly improves user personalization and accuracy.
[0081] In some embodiments, step 103, “constructing an object profile of the target object based on fused object features and historical data”, can be achieved by performing the following steps: predicting the object state of the target object based on fused object features; and constructing an object profile of the target object by combining the object state and historical data.
[0082] Here, a suitable model (such as a Transformer-based temporal decoder) is first used to process the features of the fused object to predict the object's state (i.e., the object's emotion). Taking the prediction of a player's emotion in a game as an example, the model analyzes the facial expressions, voice, and other features expressed by the fused object's features to determine whether the player is currently in an excited, anxious, or calm emotional state; it can also predict other state indicators such as the player's focus and fatigue. Then, historical data of the target object is collected. This historical data covers the target object's past object state sequences and long-term behavioral habits (such as game completion records and preferred game modes). The predicted current object state is combined with the historical data to mine the target object's behavioral patterns, object state change patterns, and other features to construct an object profile of the target object. For example, if a player has repeatedly experienced anxiety in high-difficulty levels in the past, and anxiety is also predicted in the current situation, combined with their long-term habit of exploring, an object profile of "highly exploratory and prone to anxiety in high-difficulty levels" can be constructed. This object profile can comprehensively depict the characteristics of the target object, providing a basis for subsequent personalized adjustment strategies for scene parameters (such as game difficulty adjustment, hint strategies, etc.).
[0083] By applying the above embodiments, predicting the state of an object based on its fused features allows for real-time understanding of the target object's current situation, providing an immediate basis for subsequent profile construction. Furthermore, combining historical data to build an object profile comprehensively and dynamically presents the target object's characteristics and behavioral patterns. This approach facilitates accurate object understanding, supporting personalized services and decision-making. It also allows for the discovery of changing trends and potential needs of the object through the combination of historical data and real-time status, enhancing the depth and breadth of object cognition. This helps to dynamically adjust scene parameters to better suit the actual situation of the object, increasing relevance and effectiveness.
[0084] In some embodiments, the fusion object features belong to a fusion feature sequence, which includes fusion object features for multiple time steps; see also Figure 3C "Predicting the object state of the target object based on the fused object features" can be achieved by performing the following steps 201-203: Step 201: Perform attention processing on the fused feature sequence at multiple time scales to obtain the attention feature sequence at each time scale, and concatenate the attention feature sequences at multiple time scales to obtain the multi-scale attention feature sequence; Step 202: Transform the multi-scale attention features at each time step in the multi-scale attention feature sequence to obtain local temporal features, and perform global pooling processing on the local temporal features at multiple time steps to obtain global temporal features; Step 203: Predict the object state based on the global temporal features to obtain the object state of the target object.
[0085] For step 201, the fused feature sequence includes the fused object features at each of multiple time steps. Multiple attention heads can be pre-set, each corresponding to a different time scale. Different time scales refer to different numbers of time steps included in the time scale; for example, a time step can be 1 second, and the time scale can be 3 seconds, 5 seconds, 20 seconds, etc. For each attention head, attention processing is performed on the fused feature sequence at the corresponding time scale to obtain the attention feature sequence corresponding to that time scale. This attention feature sequence includes attention features obtained by processing the fused object features at that time scale within the fused feature sequence. In this way, fused feature sequences corresponding to multiple time scales can be obtained. Continuing, the attention feature sequences at multiple time scales are concatenated to obtain a multi-scale attention feature sequence. This multi-scale attention feature sequence fuses the temporal features of multiple time scales. The structure of multiple attention heads can capture the trends and inflection points of object states within a range of several seconds to tens of seconds, and can identify short-term state transitions such as "from excitement → cooling → focus". Based on step 201, multi-timescale attention is used to simultaneously understand short-term fluctuations and long-term trends, avoiding the one-sidedness of time-series information.
[0086] For step 202, the local dynamics of each time step are first mined, and then the local dynamics of each time step are aggregated into global temporal patterns to avoid focusing only on the global picture while ignoring details, or focusing only on details while losing sight of the whole. That is, the multi-scale attention features of each time step in the multi-scale attention feature sequence are transformed to obtain local temporal features. For example, a feedforward neural network with residual connections can be used to perform a nonlinear transformation on the multi-scale attention features to obtain local temporal features, thereby enhancing the expressive power of the local temporal features. Then, global pooling (such as average pooling or global max pooling) is performed on the local temporal features of all time steps to obtain global temporal features, thus compressing the multi-scale attention feature sequence into a single global temporal feature. In practical applications, steps 201-202 can be implemented using a Transformer-based Temporal Decoder. Based on step 202, the focus is first on the local details of each time step, and then aggregated into global patterns, ensuring that the features are both detailed and comprehensive.
[0087] For step 203, the abstract global temporal features are transformed into interpretable object states (such as emotion level (such as arousal-valence two-dimensional space), attention, fatigue, cognitive load level, etc.). For example, a multilayer perceptron (MLP) can be used as a classifier for object states to achieve object state prediction. Specifically, the global temporal features are input into the multilayer perceptron, and the object state is predicted by the multilayer perceptron (such as after linear transformation + activation function (such as ReLU, Softmax)) to obtain the category or value of the object state of the target object, for example: (1) If the emotion level is predicted (such as arousal-valence two-dimensional emotion space): the MLP outputs a 2-dimensional vector, which corresponds to arousal and valence respectively, such as "emotion: arousal 7, valence 3"; (2) If the attention or fatigue is predicted: the MLP outputs a 1-dimensional value (such as attention score of 0-100, fatigue level of 0-5), such as "attention: 60; fatigue: mild". Based on step 203, the abstract features are transformed into actionable state indicators, providing a basis for decision-making in the dynamic adjustment of subsequent scene parameters (such as difficulty adjustment and UI adaptation).
[0088] By applying steps 201-203 above and employing multi-timescale attention processing, temporal relationships of different durations can be captured; multi-scale features are concatenated to enrich the information dimensions. Local temporal features are obtained by transforming multi-scale features, and then global pooling is used to extract global information, taking into account both local details and overall trends. Based on global temporal feature prediction, the object state of the target object can be grasped more accurately, improving the accuracy and comprehensiveness of object state recognition, providing a reliable basis for the dynamic adjustment of scene parameters, and optimizing user experience and decision-making efficiency.
[0089] In some embodiments, "constructing an object profile of a target object by combining object state and historical data" can be achieved through the following steps: combining object state and historical data, invoking a profile construction model to construct an object profile of the target object; based on this, see [link to relevant documentation]. Figure 3DThe portrait construction model can be constructed by executing the following steps 301-305: Step 301, extract the behavioral features of the target object from the object data, and predict the portrait label based on the behavioral features to obtain the portrait label of the target object; Step 302, obtain the portrait construction model to be updated, and predict the portrait label based on the behavioral features to obtain the predicted portrait label. The portrait construction model to be updated has a first model parameter; Step 303, update the first model parameter based on the difference between the predicted portrait label and the portrait label to obtain a second model parameter, and encrypt the second model parameter to obtain an encrypted model parameter; Step 304, upload the encrypted model parameter, wherein the encrypted model parameter is used by the server to aggregate the encrypted model parameter and the target encrypted model parameter uploaded by the target terminal to obtain an aggregated model parameter, and update the first model parameter of the portrait construction model to be updated to the aggregated model parameter to obtain the portrait construction model. The target terminal belongs to other objects participating in the virtual scene that are different from the target object; Step 305, receive the portrait construction model returned based on the encrypted model parameter, and update the portrait construction model to be updated to the portrait construction model.
[0090] Here, the construction of object profiles is based on a profile construction model. The training process of the profile construction model will be explained next.
[0091] For step 301, the object data (such as facial features, voice, physiological signals, and interaction events) can be multimodal data generated by the user participating in a virtual scene on their local device. After collecting the object data, the behavioral features of the target object (i.e., the aforementioned object features) are extracted from the object data, such as eye movement direction, tone of voice changes, and heart rate changes. After obtaining the behavioral features of the target object, a profile label prediction is performed based on the behavioral features to obtain the profile label of the target object. Specifically, based on the behavioral features of the target object, a clustering algorithm (such as K-Means algorithm or GMM algorithm) is used to predict the target object to obtain the profile label to which the target object belongs (such as aggressive player or exploratory player). The behavioral features are used as training data, and the profile labels corresponding to the behavioral features are used as labels for the training data, forming a training data pair ([behavioral feature: decreased heart rate + focused eye movement, profile label: conservative explorer]), for use in the subsequent training of the profile building model.
[0092] In step 302, the profile construction model to be updated is obtained. This model is the one that will participate in the model parameter update. The model parameters of the profile construction model to be updated are the model parameters previously synchronized from the cloud, denoted as the first model parameters. In practical applications, an initial profile construction model can be built locally using an LSTM+Transformer model, for example, based on behavioral features. Then, the initial profile construction model is updated using the model parameters synchronized from the cloud to obtain the profile construction model to be updated. After obtaining the profile construction model to be updated, profile label prediction is performed on the behavioral features based on the profile construction model to obtain the predicted profile labels.
[0093] For step 303, firstly, the difference between the predicted image label and the image label is obtained. Then, based on this difference, the first model parameters are updated to obtain the second model parameters. For example, the first model parameters are updated using algorithms such as backpropagation (e.g., fine-tuning the weights of the LSTM and / or the Transformer) to obtain the second model parameters. The second model parameters are then encrypted to obtain the encrypted model parameters.
[0094] In step 304, the encrypted model parameters are uploaded to a server (such as a cloud server). This ensures that user privacy (i.e., the original object data) is not leaked; only the encrypted model parameters are uploaded to the server. The server receives encrypted model parameters from various terminals (including the target object's terminal and target terminals of other objects participating in the virtual scene that are different from the target object). It aggregates the encrypted model parameters from all terminals (including the encrypted model parameters uploaded by the target object's terminal and the target encrypted model parameters uploaded by the target terminal) (e.g., using federated learning algorithms such as FedAvg or FedProx) to obtain aggregated model parameters. The server also includes the aforementioned profile construction model to be updated. At this point, the first model parameter of the profile construction model to be updated is updated to the aggregated model parameter to obtain the profile construction model. Then, the profile construction model is returned to each terminal, such as the target object's terminal.
[0095] In step 305, the profile construction model returned based on the encrypted model parameters is received, thereby updating the local profile construction model to be updated to the profile construction model. In this way, an object profile of the target object can be constructed based on the profile construction model for use in adjusting scene parameters.
[0096] Applying steps 301-305 above, behavioral features are extracted to predict profile labels. These labels are then combined with the prediction results from the model built based on the profile to be updated to update the model parameters, which are then encrypted and uploaded. The server aggregates encrypted model parameters from multiple terminals, updates the model built based on the profile to be updated, and returns the updated data. This process accurately generates personalized profiles, protects privacy through encryption, and utilizes federated learning aggregation to leverage the advantages of multi-terminal data while preventing data leakage. It also allows for rapid iterative model optimization, providing a more accurate profile foundation for subsequent dynamic adjustment of scene parameters and enhancing the personalized experience and adaptability of virtual scenes.
[0097] In other embodiments, object profiles can also be obtained based on rule-driven modeling, such as constructing a profile that requires reducing the game difficulty if a game fails 3 times in a row; or, a profile can be constructed by averaging the behavior of user groups, such as a novice mode profile and an expert mode profile.
[0098] Step 104: Obtain the environmental data of the target object, and based on the environmental data, predict the target scene in which the target object is located.
[0099] For step 104, environmental data of the target object is acquired, such as data collected through devices (e.g., mobile phones), ambient light sensors, and temperature sensors. This data includes, but is not limited to, illumination data (e.g., the light intensity around the device supporting the virtual scene), noise data (the noise level around the device supporting the virtual scene), and geographic location data (e.g., GPS data of the device supporting the virtual scene). Then, based on the environmental data, the target scene in which the target object is located is predicted. For example, a pre-trained classification model can be used to identify the target scene in which the target object is currently located (e.g., indoor scene, outdoor scene, cinema scene, nighttime scene, conference scene, outdoor bright light + high noise scene, indoor low light + quiet scene, etc.) based on the environmental data.
[0100] Step 105: Adjust the scene parameters of the virtual scene based on the object profile and the target scene.
[0101] In step 105, after obtaining the object profile and target scene, the scene parameters of the virtual scene are adjusted based on the object profile and target scene. Thus, through multimodal data fusion and historical data modeling, an accurate object profile is constructed, and the target scene where the object is located is predicted using environmental data. Therefore, based on the object profile and target scene, the virtual scene parameters are dynamically adjusted, significantly improving the adaptability of the virtual scene to the individual user and enhancing the user experience and immersion.
[0102] In some embodiments, step 105, "adjusting the scene parameters of the virtual scene based on the object profile and the target scene," can be achieved by performing the following steps: adjusting the first scene parameters of the virtual scene based on the object profile, and adjusting the second scene parameters of the virtual scene based on the target scene; wherein the scene parameters include the first scene parameters and the second scene parameters, and the scene parameters include at least one of the following: the character parameters of the player character, the character parameters of the non-player character, the interface parameters of the human-computer interaction interface of the virtual scene, the audio parameters of the virtual scene, the visual parameters of the virtual scene, the scene guidance parameters of the virtual scene, and the interaction parameters of the virtual scene.
[0103] Here, the first scene parameters of the virtual scene are adjusted based on the object profile, and then the second scene parameters of the virtual scene are adjusted based on the target scene. The first scene parameters and the second scene parameters may overlap (i.e., include the same scene parameters) or they may not overlap (i.e., do not include the same scene parameters). Scene parameters include at least one of the following: player character parameters (such as the density and number of enemy players, the attack power, movement speed, health, skill strength, etc. of friendly players), non-player character parameters (such as the behavior strategies of NPCs and AI characters, such as attack frequency, skill release probability, attack range, etc.), interface parameters of the human-computer interaction interface of the virtual scene (such as the number and size of interface buttons, the complexity of UI operations, interface transparency, information display density, etc.), audio parameters of the virtual scene (such as music rhythm, ambient sound volume, sound effect tension), visual parameters of the virtual scene (such as visual color, brightness, style, complexity of visual effects), scene guidance parameters of the virtual scene (such as the display frequency, display position, prominence, and display timing of guidance and prompt information), interaction parameters of the virtual scene (such as game difficulty, interaction cooldown time, interaction success rate, skill cooldown time, game resource refresh frequency, game resource acquisition amount, game scene complexity, game scene interference factors (such as severe weather, terrain traps, etc.), etc.), and device performance adaptation parameters (such as graphics quality (texture resolution, model accuracy), frame rate limit).
[0104] By applying the above embodiments, the rich and multi-dimensional scene parameters can be dynamically adjusted based on user profiles and scenarios, which can accurately adapt to users and allow different users to have an immersive and comfortable virtual scene experience in various scenarios, thereby improving retention and satisfaction. Furthermore, it increases the diversity of dynamically adjusted scene parameters, making the adjustment effect of scene parameters better and more in line with the diverse needs of users.
[0105] In some embodiments, "adjusting the first scene parameter of the virtual scene based on the object profile" can be achieved by performing the following steps: obtaining the target mapping relationship between the target parameter values of the object profile and the first scene parameter; adjusting the parameter value of the first scene parameter from the current parameter value to the target parameter value according to the target mapping relationship; based on this, see... Figure 3E For each object sample among multiple object samples, the following steps 401-403 are performed respectively to obtain the mapping relationship corresponding to the object sample: Step 401, obtain the behavioral features of the object sample, and predict the portrait label based on the behavioral features to obtain the portrait label of the object sample; Step 402, obtain the scene parameter value corresponding to the portrait label; Step 403, generate the portrait sample of the object sample based on the behavioral features, and fit the mapping relationship between the portrait sample and the scene parameter value to obtain the mapping relationship corresponding to the object sample; wherein, the mapping relationship corresponding to multiple object samples includes the target mapping relationship.
[0106] Here, a mapping relationship is pre-built between portrait samples and scene parameter values. The portrait sample includes the object portrait, the scene parameter includes the target parameter value of the first scene parameter, and the mapping relationship includes the target mapping relationship. There are multiple mapping relationships, with a one-to-one correspondence between portrait samples and scene parameter values. For example, the mapping relationship could include: portrait sample "High Anxiety + Novice Player" → scene parameter value "Game Difficulty: Easy"; portrait sample "Late Night + Low Brightness Environment" → "UI Brightness: 80% (High Brightness)"; portrait sample "Quiet Scene" → scene parameter value "Sound Effect Intensity: High". Based on this, we can first find the target mapping relationship between the object portrait and the target scene parameter value of the first scene parameter from multiple mapping relationships. Then, according to the target mapping relationship, the parameter value of the first scene parameter is adjusted from the current parameter value to the target parameter value. For example, if the object portrait is "High Anxiety + Novice Player", then the parameter value of the scene parameter "Game Difficulty" can be adjusted from the current parameter value "Hard" to the target parameter value "Easy". The following explains the process of building the mapping relationship.
[0107] Multiple object samples were collected to construct mapping relationships. Simultaneously, object data from multiple modalities (e.g., visual, speech, physiological, and environmental modalities) were also collected for each object sample. Object features were then extracted from the object data of each modality, and these multi-modal object features were fused to obtain the fused object features of the object sample. These fused object features are then used as the behavioral features of the object sample. In practical applications, these behavioral features also include historical behavioral features extracted from the object sample's historical data (e.g., historical operation frequency, dwell time, number of failures, etc.). The following processing is then performed on each object sample:
[0108] For step 401, after obtaining the behavioral features of the object sample, a profile label prediction is performed based on the behavioral features to obtain the profile label of the object sample. Specifically, based on the behavioral features of the object sample, a clustering algorithm (such as K-Means algorithm or GMM algorithm) is used to predict the object sample to obtain the profile label to which the object sample belongs (such as aggressive player or exploratory player).
[0109] For step 402, multiple candidate profile labels are pre-set, and corresponding scene parameter values (which can also be understood as scene parameter adjustment strategies) are set for each candidate profile label. These multiple candidate profile labels include the profile label to which the object sample belongs. For example, the candidate profile label "high anxiety" corresponds to the scene parameter value "game difficulty = medium to low difficulty"; the candidate profile label "quiet scene" corresponds to the scene parameter value "audio intensity = silent". Based on this, the profile label is found from the multiple candidate profile labels, thereby obtaining the scene parameter value corresponding to the profile label.
[0110] For step 403, firstly, a portrait sample of the object sample is generated based on behavioral features. For example, a portrait model can be used to construct the portrait sample based on behavioral features (specifically, the portrait vector of the portrait sample). Then, a mapping relationship (or functional relationship) between the portrait sample and the scene parameter values is fitted. For example, a multinomial fitting or a neural network (MLP) is used to fit the mapping relationship between the portrait sample and the scene parameter values. Taking the scene parameter value as the game difficulty as an example, D... t =f(P t ), where D t For game difficulty (e.g., easy), P t Let f be the image sample, and f be the mapping relationship obtained by fitting.
[0111] By applying steps 401-403 above, and by collecting behavioral features from multiple dimensions and generating precise profile tags, personalized user differentiation can be achieved; by matching verified scene parameters, the basic effectiveness of the adjustment strategy can be ensured; and by fitting the dynamic mapping relationship between the profile samples and scene parameters, the reliability of the adjustment strategy can be ensured by relying on the commonalities of the group, while also achieving precise adjustment for individual behaviors, realizing a personalized interactive experience, and optimizing the user's immersion and retention rate in the virtual scene.
[0112] In some embodiments, the following steps may also be performed: determining the target state deviation between the object state and the target object state; obtaining the state offset weight and state change control factor of the object state, and determining the total state deviation within a set time period before the current time point based on the time decay factor; wherein, the state change control factor is used to control the change trend of the state deviation at future time points after the current time point so that the change trend is in a smooth state; the time decay factor is used to control the deviation weight of the state deviation at historical time points within the set time period so that the deviation weight is negatively correlated with the time interval between the historical time point and the current time point; determining the parameter adjustment value of the scene parameters based on the target state deviation, state offset weight, total state deviation, and state change control factor; based on this, step 105 "adjusting the scene parameters of the virtual scene based on the object profile and the target scene" can be achieved by performing the following steps: adjusting the scene parameters of the virtual scene according to the parameter adjustment value based on the object profile and the target scene.
[0113] Here, the target object state is a pre-set desired state that the object should maintain in the virtual scene, such as a state of moderate excitement and low anxiety, which might correspond to an excitement value of 0.6 and an anxiety value of 0.2. Therefore, we first determine the target state deviation between the current object state and the target object state, i.e., target state deviation e(t) = target object state E_target - object state E(t). Then, we obtain the state offset weight P and the state change control factor D of the object state. The state offset weight P is an immediate response to the target state deviation, used to control the adjustment of scene parameters in response to the target state deviation. The larger the target state deviation, the larger the state offset weight, and the stronger the adjustment, thus quickly correcting the current target state deviation by adjusting scene parameters. The state change control factor D controls the trend of the state deviation (i.e., the deviation between the current object state and the target object state) at future time points after the current time point, ensuring that the trend is smooth (i.e., the rate of change of the state deviation indicated by the trend is below the rate threshold). When the rate of change of the state deviation exceeds the rate threshold, the state change control factor decreases to reduce the adjustment of scene parameters, thus preventing the state deviation from reversing due to excessive adjustment of scene parameters. Simultaneously, based on the time decay factor, the total state deviation I within a set time period before the current time point is determined. The time decay factor controls the deviation weight of the state deviation at historical time points within the set time period, ensuring that the deviation weight is negatively correlated with the time interval between the historical time point and the current time point. The set time period can be a time period adjacent to the current time point, meaning the end time of the set time period is adjacent to the current time point. The total state deviation is the accumulation of state deviations over a set time period. By accumulating state deviations over a period of time, it ensures that the state deviation can eventually be eliminated by adjusting scene parameters. The earlier the time point (i.e., the larger the time interval between the historical time point and the current time point), the smaller the deviation weight of the state deviation, to avoid the old deviation excessively affecting the adjustment of the current scene parameters. Finally, based on the target state deviation, state offset weight, total state deviation, and state change control factor, combined with a preset adjustment mapping relationship, the parameter adjustment value of the scene parameters is determined (e.g., game parameter +0.1). This adjustment mapping relationship refers to the mapping relationship between the parameter adjustment value of the scene parameters and the target state deviation, state offset weight, total state deviation, and state change control factor. Therefore, when adjusting scene parameters, the scene parameters of the virtual scene are adjusted according to the parameter adjustment value.
[0114] See Figure 5The process of determining the parameter adjustment value includes: 1. Obtaining the current object state (i.e., estimated emotion) E(t); 2. Obtaining the target object state (i.e., target emotion) E_target; 3. Calculating the target state deviation e(t) = E_target - E(t); 4. Determining the parameter adjustment value (e.g., game difficulty + 0.1) through a PID-like controller, where P is the state offset weight, I is the total state deviation within a set time period, and D is the state change control factor. Here, a PID-like closed-loop algorithm is used to gradually and stably adjust the deviation of the object state.
[0115] By applying the above embodiments, scene parameters can be dynamically and accurately adjusted by determining the target state deviation and obtaining the state offset weight. The state offset weight can quickly respond to the current target state deviation, the total state deviation (including time decay) prevents small deviations from persisting for a long time, the state change control factor smooths the adjustment trend, and the time decay factor makes the impact of historical deviations reasonable. The combination of multiple factors ensures both timely adjustment and avoids over- or under-adjustment, making scene parameter adaptation more accurate and stable, and improving user experience and scene adaptability.
[0116] In other embodiments, in addition to adjusting scene parameters based on the above-mentioned object profile and target scene, scene parameters that support user self-selection can also be provided (such as day mode, night mode, travel mode, game difficulty: easy, normal, hard) so that users can manually select the required scene parameters as needed; or the automatic switching between night mode and day mode can be realized based on the device's system clock.
[0117] In practical applications, scene parameters can be adjusted according to adjustment strategies. Adjustment strategies are used to indicate how to adjust scene parameters, and can include the above-mentioned parameter adjustment values, mapping relationships, etc. For example, adjustment strategies can include: adjusting the parameter value of the scene parameter "game difficulty" from the current parameter value "hard" to the target parameter value "easy"; adjusting the parameter value of the scene parameter "sound effect intensity" from the current parameter value "high" to the target parameter value "low"; increasing the parameter value of the scene parameter "game difficulty" by 0.1 from the current parameter value. After the server generates the adjustment strategy, the adjustment strategy can be distributed to the terminal in the following ways: (1) Distribution timing: a) Real-time push: The adjustment strategy is pushed as soon as the object state changes to trigger scene parameter updates. b) Periodic push: The adjustment strategy is pushed according to the set period, such as pushing an adjustment strategy once at the beginning of each session (such as a complete game start-up to exit) to update the scene parameters once. c) Event trigger: The adjustment strategy is pushed when a set event occurs to trigger scene parameter updates, such as a player consecutive failure event or a game interruption event. (2) Distribution Granularity: a) Global Distribution: Push adjustment strategies for all scene parameters matching the current object profile and / or target scene at once to update all scene parameters matching the current object profile and / or target scene at once. b) Module Distribution: Push adjustment strategies for scene parameters related to the current module (e.g., sound effect cues intensity) only to adjust scene parameters related to the current module. c) Hierarchical Distribution: Set different levels and different distribution methods for different levels. For example, set different distribution strategies for different levels (e.g., different device levels, player groups). Level 1 (high-performance devices or core players) uses a highly dynamic distribution method (the push frequency of adjustment strategies is greater than or equal to the frequency threshold), while Level 1 (low-performance devices or novice players) uses a conservative distribution method (the push frequency of adjustment strategies is less than the frequency threshold). 3) Delayed Loading of Adjustment Strategies: Avoid performance fluctuations caused by frequent adjustment of scene parameters.
[0118] In practical applications, to ensure data security, data can be classified for privacy. Specifically, data with a sensitivity level below the threshold (such as playtime) can be sent to the server (or cloud) for processing; data with a sensitivity level greater than or equal to the threshold (such as facial emotions) can be processed locally or sent to the server (or cloud) after differential privacy processing.
[0119] In practical applications, this application embodiment also has high-performance scalable deployment and multi-dimensional adjustment capabilities. Specifically, (1) Resource elastic management: The cloud uses container cloud or Kubernetes to support the elastic expansion of services, dynamically allocates resources according to concurrency and controls costs. (2) Multi-dimensional control interface: Defines a unified API (REST / GRPC), which can modify multi-dimensional scene parameters (such as the density of adversaries, UI logic complexity, music mixing parameters, visual style and prompt intensity, etc.), supports flexible invocation of multiple virtual scenes (such as games), supports deployment on general platforms and has good cross-game compatibility. (3) Performance guarantee strategy: The terminal is only responsible for necessary sensor collection and local terminal caching, and the actual calculation is transferred to the cloud to ensure that the frame rate and visual quality of the virtual scene on the terminal are not affected.
[0120] In practical applications, the cloud computing power and closed-loop feedback control in this embodiment are integrated, specifically: (1) Architecture design: a hybrid cloud architecture (Edge+Central) is adopted, and lightweight models are deployed at the edge / local for preprocessing (such as the collection of object data of multiple modalities). Complex reasoning (such as the prediction of object state and the construction of object profile) and model optimization are completed by the cloud, and the local model (such as the profile construction model) is continuously calibrated to reduce hardware requirements and user-end computing burden and running pressure; and the existing mobile device capabilities can support it, without relying on external hardware, supporting deployment on general platforms and good cross-game compatibility. (2) Feedback control module: a PID-like closed-loop algorithm with time decay factor and smoothing coefficient (i.e. state change control factor) is designed to determine the parameter adjustment value, thereby adjusting the scene parameters based on the parameter adjustment value to gradually and smoothly adjust the deviation of the object state, avoid abrupt behavior, and save the adjustment history and effect in the cloud for use in cloud reinforcement learning to optimize future decisions. (3) Online training and inference: The model is updated regularly in the cloud (such as the model for predicting the state of an object, the model for building a profile, etc.), and the reinforcement learning strategy is optimized through the replay buffer, shortening the model update cycle to the minute level.
[0121] In this embodiment, each type of object profile is assigned an appropriate mapping relationship (e.g., exploratory profiles prefer low visual interference and high cue frequency). The scene parameters of each dimension are dynamically adjusted in combination with the real-time object status and scene inference results (i.e., target scene). The parameter adjustment values of the scene parameters are based on the output of the closed-loop feedback controller (PID controller). The interface uniformly drives the linkage adjustment of scene parameters of multiple dimensions, avoiding the adjustment conflict of scene parameters and giving the overall adjustment a more flexible and comprehensive effect.
[0122] Applying the embodiments described above, firstly, object data of the target object participating in the virtual scene in multiple modalities is acquired, and object features of the object data in each modality are extracted. Then, the object features of multiple modalities are fused to obtain fused object features. Secondly, historical data of the target object's participation in the virtual scene is acquired, and an object profile of the target object is constructed based on the fused object features and historical data. Thirdly, environmental data of the target object is acquired, and the target scene in which the target object is located is predicted based on the environmental data. Finally, the scene parameters of the virtual scene are adjusted by combining the object profile and the target scene. Thus, by using multimodal data fusion and historical data modeling to accurately construct an object profile, and by predicting the target scene in which the object is located using environmental data, the virtual scene parameters can be dynamically adjusted based on the object profile and the target scene. This significantly improves the adaptability of the virtual scene to the individual user, enhancing the user experience and immersion.
[0123] The following description continues to illustrate the exemplary structure of the virtual scene data processing device 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules in the virtual scene data processing device 555 stored in the memory 550 may include: an extraction module 5551, used to acquire object data of multiple modalities of a target object participating in the virtual scene, and extract object features of the object data of each modality; a fusion module 5552, used to fuse the object features of multiple modalities to obtain fused object features; a construction module 5553, used to acquire historical data of the target object participating in the virtual scene, and construct an object profile of the target object based on the fused object features and the historical data; a prediction module 5554, used to acquire environmental data of the target object, and predict the target scene in which the target object is located based on the environmental data; and an adjustment module 5555, used to adjust the scene parameters of the virtual scene by combining the object profile and the target scene.
[0124] In some embodiments, the extraction module 5551 is further configured to: if the modality is a visual modality and the object data includes an object image sequence of the target object, extract local image features of each object image in the object image sequence, and extract temporal image features between the local image features of the object image sequence, using the temporal image features as the object features; if the modality is a speech modality and the object data includes object speech of the target object, extract emotional frequency band features, prosodic features, and spectral features of the object speech, and use the emotional frequency band features, the prosodic features, and the spectral features as... The object features are as follows: If the modality is a physiological modality and the object data includes the physiological data sequence of the target object, local physiological features of the physiological data sequence are extracted in each time window, and physiological temporal features between the local physiological features of the physiological data sequence are extracted, and the physiological temporal features are used as the object features; if the modality is an environmental modality and the object data includes the environmental data sequence of the target object, the environmental data sequence is subjected to temporal embedding processing to obtain environmental temporal features, and the environmental temporal features are mapped to environmental label features, and the environmental label features are used as the object features.
[0125] In some embodiments, the fusion module 5552 is further configured to, for each of the plurality of modalities, perform attention processing using the object features of the first modality as a query vector and the object features of at least one second modality as a key vector and a value vector, to obtain attention weights between each second modality and the first modality, wherein the first modality is any of the plurality of modalities and the second modality is a modality different from the first modality among the plurality of modalities; for each first modality, perform weighted fusion of the object features of at least one second modality according to the attention weights of each second modality to obtain a first fused feature, and fuse the first fused feature and the object features of the first modality to obtain a second fused feature of the first modality; and map the second fused feature of each first modality to a shared semantic space to obtain the fused object feature.
[0126] In some embodiments, the construction module 5553 is further configured to predict the object state of the target object based on the fused object features; and to construct an object profile of the target object by combining the object state and the historical data.
[0127] In some embodiments, the fused object features belong to a fused feature sequence, which includes the fused object features at multiple time steps. The construction module 5553 is further configured to perform attention processing on the fused feature sequence at multiple time scales to obtain attention feature sequences at each time scale, and to concatenate the attention feature sequences at multiple time scales to obtain a multi-scale attention feature sequence. The multi-scale attention features at each time step in the multi-scale attention feature sequence are transformed to obtain local temporal features, and the local temporal features at each of the multiple time steps are subjected to global pooling processing to obtain global temporal features. Based on the global temporal features, object state prediction is performed to obtain the object state of the target object.
[0128] In some embodiments, the construction module 5553 is further configured to combine the object state and the historical data, invoke a profile construction model, and construct an object profile of the target object; the construction module 5553 is further configured to extract behavioral features of the target object from the object data, and perform profile label prediction based on the behavioral features to obtain a profile label of the target object; obtain a profile construction model to be updated, and perform profile label prediction based on the behavioral features to obtain a predicted profile label, wherein the profile construction model to be updated has first model parameters; and adjust the first model parameters based on the difference between the predicted profile label and the profile label. The system obtains second model parameters and encrypts them to obtain encrypted model parameters. It then uploads these encrypted model parameters, which are used by the server to aggregate the encrypted model parameters and the target encrypted model parameters uploaded by the target terminal to obtain aggregated model parameters. The server then updates the first model parameters of the portrait construction model to be updated to the aggregated model parameters, resulting in the portrait construction model. The target terminal belongs to another object participating in the virtual scene that is different from the target object. Finally, the system receives the portrait construction model returned based on the encrypted model parameters and updates the portrait construction model to be updated to the original portrait construction model.
[0129] In some embodiments, the adjustment module 5555 is further configured to: determine the target state deviation between the object state and the target object state; obtain the state offset weight and state change control factor of the object state; and determine the total state deviation within a set time period before the current time point based on a time decay factor; wherein, the state change control factor is used to control the change trend of the state deviation at future time points after the current time point, so that the change trend is in a smooth state; the time decay factor is used to control the deviation weight of the state deviation at historical time points within the set time period, so that the deviation weight is negatively correlated with the time interval between the historical time point and the current time point; determine the parameter adjustment value of the scene parameters based on the target state deviation, the state offset weight, the total state deviation, and the state change control factor; the adjustment module 5555 is further configured to: adjust the scene parameters of the virtual scene according to the parameter adjustment value based on the object profile and the target scene.
[0130] In some embodiments, the adjustment module 5555 is further configured to adjust a first scene parameter of the virtual scene based on the object profile, and to adjust a second scene parameter of the virtual scene based on the target scene; wherein the scene parameters include the first scene parameter and the second scene parameter, and the scene parameters include at least one of the following: character parameters of the player character, character parameters of the non-player character, interface parameters of the human-computer interaction interface of the virtual scene, audio parameters of the virtual scene, visual parameters of the virtual scene, scene guidance parameters of the virtual scene, interaction parameters of the virtual scene, and device performance adaptation parameters.
[0131] In some embodiments, the adjustment module 5555 is further configured to obtain a target mapping relationship between the object profile and the target parameter value of the first scene parameter; adjust the parameter value of the first scene parameter from the current parameter value to the target parameter value according to the target mapping relationship; the adjustment module 5555 is further configured to perform the following processing for each of the multiple object samples: obtain the behavioral features of the object sample, and perform profile label prediction based on the behavioral features to obtain the profile label of the object sample; obtain the scene parameter value corresponding to the profile label; generate a profile sample of the object sample based on the behavioral features, and fit the mapping relationship between the profile sample and the scene parameter value to obtain the mapping relationship corresponding to the object sample; wherein, the mapping relationship corresponding to the multiple object samples includes the target mapping relationship.
[0132] It should be noted that the description of the device embodiments in this application is similar to the description of the method embodiments described above, and has similar beneficial effects as the method embodiments, so it will not be repeated here. Any technical details not covered in the data processing device for virtual scenes provided in the embodiments of this application can be understood based on the description of the technical details in the above method embodiments.
[0133] This application also provides a computer program product, which includes computer-executable instructions or a computer program stored in a computer-readable storage medium. The processor of an electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, causing the electronic device to perform the data processing method for a virtual scene provided in this application.
[0134] This application also provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the data processing method for the virtual scene provided in this application.
[0135] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0136] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0137] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0138] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0139] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A data processing method for a virtual scene, characterized in that, The method includes: Obtain object data of multiple modalities of the target object participating in the virtual scene, and extract object features of the object data of each modality; The object features of multiple modalities are fused to obtain fused object features; The process involves acquiring historical data of the target object's participation in the virtual scene, and constructing an object profile of the target object based on the fused object features and the historical data. Specifically, the object state of the target object is predicted based on the fused object features. The object profile of the target object is then constructed by combining the object state and the historical data. The historical data includes behavioral trajectory data of the target object participating in the virtual scene, which includes game interaction data and state time-series data. The game interaction data includes completion records, number of failures, and operational behaviors. The state time-series data includes emotional curves and physiological state time sequences. The environmental data of the target object is obtained, and the target scene in which the target object is located is predicted based on the environmental data. Determine the target state deviation between the stated object state and the target object state; Obtain the state offset weight and state change control factor of the object state, and determine the total state deviation within a set time period before the current time point based on the time decay factor; The state change control factor is used to control the trend of state deviation changes at future time points after the current time point so that the trend of change is in a smooth state. When the rate of change of state deviation is greater than the rate threshold, the state change control factor will be reduced to reduce the adjustment intensity of scene parameters. The time decay factor is used to control the deviation weight of the state deviation of historical time points within the set time period, so that the deviation weight is negatively correlated with the time interval between the historical time point and the current time point. Based on the target state deviation, the state offset weight, the total state deviation, and the state change control factor, the parameter adjustment values of the scene parameters of the virtual scene are determined; Based on the object profile and the target scene, the scene parameters of the virtual scene are adjusted according to the parameter adjustment values.
2. The method as described in claim 1, characterized in that, The object features extracted from the object data of each modality include: If the modality is a visual modality and the object data includes an object image sequence of the target object, extract the local image features of each object image in the object image sequence, and extract the temporal image features between the local image features of the object image sequence, and use the temporal image features as the object features; If the modality is a speech modality and the object data includes the object speech of the target object, extract the emotional frequency band features, prosodic features and spectral features of the object speech, and use the emotional frequency band features, the prosodic features and the spectral features as the object features; If the modality is a physiological modality and the object data includes the physiological data sequence of the target object, extract the local physiological features of the physiological data sequence in each time window, and extract the physiological temporal features between the local physiological features of the physiological data sequence, and use the physiological temporal features as the object features; If the modality is an environment modality and the object data includes an environment data sequence of the target object, the environment data sequence is subjected to time embedding processing to obtain environmental temporal features, and the environmental temporal features are mapped to environmental label features, and the environmental label features are used as the object features.
3. The method as described in claim 1, characterized in that, The process of fusing the object features of multiple modalities to obtain fused object features includes: For each first modality among the plurality of modalities, attention processing is performed using the object features of the first modality as a query vector and the object features of at least one second modality as a key vector and a value vector to obtain the attention weight between each second modality and the first modality. The first modality is any one of the plurality of modalities, and the second modality is a modality among the plurality of modalities that is different from the first modality. For each first modality, the object features of at least one second modality are weighted and fused according to the attention weight of each second modality to obtain a first fused feature, and the first fused feature and the object features of the first modality are fused to obtain a second fused feature of the first modality; The second fusion feature of each of the first modalities is mapped to a shared semantic space to obtain the fusion object features.
4. The method as described in claim 1, characterized in that, The fused object features belong to a fused feature sequence, which includes the fused object features at multiple time steps; predicting the object state of the target object based on the fused object features includes: The fused feature sequence is subjected to attention processing at multiple time scales to obtain attention feature sequences at each time scale, and the attention feature sequences at multiple time scales are concatenated to obtain a multi-scale attention feature sequence. The multi-scale attention features at each time step in the multi-scale attention feature sequence are transformed to obtain local temporal features, and the local temporal features at each of the multiple time steps are subjected to global pooling to obtain global temporal features. Based on the global temporal features, the object state is predicted to obtain the object state of the target object.
5. The method as described in claim 1, characterized in that, The step of constructing an object profile of the target object by combining the object state and the historical data includes: By combining the object's state and the historical data, the profile building model is invoked to construct an object profile of the target object; The method further includes: The behavioral features of the target object are extracted from the object data, and the profile label is predicted based on the behavioral features to obtain the profile label of the target object; Obtain the profile construction model to be updated, and predict the profile label based on the behavioral features based on the profile construction model to be updated to obtain the predicted profile label. The profile construction model to be updated has a first model parameter. Based on the difference between the predicted image label and the image label, the first model parameters are updated to obtain the second model parameters, and the second model parameters are encrypted to obtain the encrypted model parameters; The encrypted model parameters are uploaded, wherein the encrypted model parameters are used by the server to aggregate the encrypted model parameters and the target encrypted model parameters uploaded by the target terminal to obtain aggregated model parameters, and the first model parameter of the profile construction model to be updated is updated to the aggregated model parameters to obtain the profile construction model, wherein the target terminal belongs to other objects participating in the virtual scene and different from the target object; Receive the portrait construction model returned based on the encrypted model parameters, and update the portrait construction model to be updated to the portrait construction model.
6. The method as described in claim 1, characterized in that, The step of adjusting the scene parameters of the virtual scene according to the parameter adjustment value based on the object profile and the target scene includes: Based on the object profile, the first scene parameters of the virtual scene are adjusted, and based on the target scene, the second scene parameters of the virtual scene are adjusted. The scene parameters include the first scene parameters and the second scene parameters, and the scene parameters include at least one of the following: character parameters of the player character, character parameters of the non-player character, interface parameters of the human-computer interaction interface of the virtual scene, audio parameters of the virtual scene, visual parameters of the virtual scene, scene guidance parameters of the virtual scene, interaction parameters of the virtual scene, and device performance adaptation parameters.
7. The method as described in claim 6, characterized in that, The step of adjusting the first scene parameters of the virtual scene based on the object profile includes: Obtain the target mapping relationship between the object profile and the target parameter values of the first scene parameters; According to the target mapping relationship, the parameter value of the first scene parameter is adjusted from the current parameter value to the target parameter value; The method further includes: For each of the multiple object samples, perform the following processing: The behavioral features of the object sample are obtained, and the image label is predicted based on the behavioral features to obtain the image label of the object sample. Obtain the scene parameter values corresponding to the portrait tags; Based on the behavioral features, a portrait sample of the object sample is generated, and a mapping relationship between the portrait sample and the scene parameter values is fitted to obtain the mapping relationship corresponding to the object sample; The mapping relationship corresponding to the multiple object samples includes the target mapping relationship.
8. A data processing device for a virtual scene, characterized in that, The device includes: An extraction module is used to acquire object data of multiple modalities of target objects participating in a virtual scene, and to extract object features of the object data of each modality; A fusion module is used to fuse the object features of multiple modalities to obtain fused object features; A construction module is used to acquire historical data of the target object's participation in the virtual scene, and to construct an object profile of the target object based on the fused object features and the historical data; wherein, based on the fused object features, the object state of the target object is predicted; the object profile of the target object is constructed by combining the object state and the historical data; the historical data includes behavioral trajectory data of the target object participating in the virtual scene, the behavioral trajectory data includes game interaction data and state time series data, the game interaction data includes level completion records, number of failures and operation behaviors, and the state time series data includes emotion curves and physiological state time series; The prediction module is used to acquire environmental data of the target object and, based on the environmental data, predict the target scene in which the target object is located. The adjustment module is used to determine the target state deviation between the object state and the target object state; obtain the state offset weight and state change control factor of the object state; and determine the total state deviation within a set time period before the current time point based on the time decay factor. The state change control factor is used to control the trend of state deviation changes at future time points after the current time point so that the trend of change is in a smooth state. When the rate of change of state deviation is greater than the rate threshold, the state change control factor will be reduced to reduce the adjustment intensity of scene parameters. The time decay factor is used to control the deviation weight of the state deviation of historical time points within the set time period, so that the deviation weight is negatively correlated with the time interval between the historical time point and the current time point. Based on the target state deviation, the state offset weight, the total state deviation, and the state change control factor, the parameter adjustment values of the scene parameters of the virtual scene are determined; The adjustment module is also used to adjust the scene parameters of the virtual scene according to the parameter adjustment value based on the object profile and the target scene.
9. An electronic device, characterized in that, include: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the data processing method for the virtual scene as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The system stores computer-executable instructions, which, when executed by a processor, implement a data processing method for a virtual scene as described in any one of claims 1 to 7.
11. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, they implement the method as described in any one of claims 1 to 7.