AI role interaction method, interaction device and interaction equipment in game platform
By obtaining game scenes and user interaction data to generate dynamic interaction strategies, adjusting AI character behavior parameters, and rendering multimodal interactive content, the problem of single AI character interaction form in the existing technology is solved, and the game's immersion and user satisfaction are improved.
Patent Information
- Application Number
- CN202510750926.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The interactive content of AI character interaction methods in existing games is single in form, the behavior is rigid and rigid, and the interaction with users lacks a sense of reality and immersion. Especially when the multimodal interactive content is smoothly transmitted and presented, it aggravates the sense of separation of the interactive experience.
By obtaining game scenes, task information and character information, generating a set of behavior parameters of AI characters, collecting user operation behavior and voice interaction data, performing intention recognition processing, generating user behavior feature vectors, determining their relationship links with game scenes, task information and character information, generating dynamic interaction strategies, adjusting behavior parameter sets, rendering multi-modal interaction content and displaying them to the user terminal.
It realizes that the interaction of AI characters is more realistic and natural, significantly improves the user's game immersion and satisfaction, and solves the problem of smooth transmission and presentation of multimodal interactive content.
Smart Images

Figure CN120381658A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of game interaction, and particularly to an AI character interaction method, interaction device, and interaction equipment within a game platform. Background Art
[0002] In existing games, the forms of interaction content presented by AI character interaction methods are single, and the behavioral performances are rigid and stereotyped. The interaction with users lacks a sense of reality and immersion, seriously affecting the game experience. Especially when smoothly transmitting and presenting multi-modal interaction content, the sense of fragmentation of the interaction experience is aggravated. Summary of the Invention
[0003] The main purpose of the present invention is to provide an AI character interaction method within a game platform, aiming to make the interaction of AI characters more real and natural, and significantly improve the game immersion and satisfaction of users.
[0004] To achieve the above purpose, the present invention provides an AI character interaction method within a game platform. The AI character interaction method includes: Obtain the current game scene, game task information, and game character information, and generate a first set of behavior parameters for the AI character; Collect the current operation behavior data and voice interaction data of the user; Perform intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, where the user behavior feature vector includes the user's core intention, operation style features, and emotional tendency; Determine the relationship link between the user behavior feature vector and at least one of the game scene, the task information, and the character information to generate a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment range; Adjust the first set of behavior parameters according to the dynamic interaction strategy to generate a second set of behavior parameters. The first set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient; Render and generate multi-modal interaction content matching the second set of behavior parameters. The multi-modal interaction content includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient; Display the multi-modal interaction content to the user terminal.
[0005] Optionally, the performing intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector includes: Perform acoustic feature standardization processing on the voice interaction data and then input it into a speech recognition engine to convert the processed voice interaction data into a text statement sequence with timestamp marks; Perform semantic segmentation of the operation behavior data based on the hot zone distribution, frequency intensity grading, and validity marking of the task progress status of the display interface of the user terminal to generate a corresponding sequence of operation action statements; Construct a spatio-temporal alignment model for the text statement sequence and the operation action statement sequence, perform timestamp dynamic matching, game object association mapping, and logical consistency detection, and generate an intermediate feature set including semantic keyword weight distribution, decision preference intensity index, and sentiment tendency index; Perform intention inference on the intermediate feature set based on the dynamically updated context state matrix, determine the intention probability distribution, and generate a user behavior feature vector corresponding to the intention probability distribution. The context state matrix includes the current task phase target, user historical behavior statistical features, and a scene object attribute relationship graph.
[0006] Optionally, the rendering and generating of multi-modal interaction content matching the second set of behavior parameters includes: Based on the dialogue response mode in the second set of behavior parameters, call a speech synthesis engine to generate speech waveform data with emotional prosody features; According to the task guidance intensity value in the second set of behavior parameters, extract basic action primitives from a preset action library, and through a physical simulation engine, mix and optimize the basic action primitives to generate three-dimensional character action data matching the task information, and generate an interface highlighting special effect positively correlated with the guidance intensity value; Based on the emotional expression coefficient in the second set of behavior parameters, generate a facial micro-expression fusion weight, and control a particle special effect engine to generate an environment feedback special effect matching the emotional dimension based on the facial micro-expression fusion weight; Input the speech waveform data, the three-dimensional character action data, the interface highlighting special effect, and the environment feedback special effect into a multi-modal synchronization engine to generate multi-modal interaction content with spatio-temporal feature alignment.
[0007] Optionally, the inputting the speech waveform data, the three-dimensional character action data, the interface highlighting special effect, and the environment feedback special effect into a multi-modal synchronization engine to generate multi-modal interaction content with spatio-temporal feature alignment includes: According to the timing requirements of the dialogue response mode, add a first-level timestamp mark to the speech waveform data; Based on the physical simulation logic of the task guidance intensity, add a second-level timestamp mark to the three-dimensional character action data; According to the dynamic change curve of the emotional expression coefficient, add a third-level timestamp mark to the environment feedback special effect; Align the first-level, second-level, and third-level timestamp marks through a dynamic offset compensation algorithm to generate a unified reference time axis; According to the unified reference timeline, input the speech waveform data, the three-dimensional character action data, the interface highlighting special effects, and the environmental feedback special effects into a multi-modal synchronization engine for spatio-temporal feature alignment processing.
[0008] Optionally, determining the relationship link between the user behavior feature vector and at least one of the game scenario, the task information, and the character information to generate a dynamic interaction strategy with interaction strategy priorities and parameter adjustment amplitudes includes: Construct a dynamic relationship graph including game scenario objects, task target nodes, and character attribute nodes; Calculate the multi-dimensional association strength between the user behavior feature vector and at least one type of node in the dynamic relationship graph; Generate an initial interaction strategy set based on the multi-dimensional association strength; Optimize the initial interaction strategy set through a strategy conflict resolution model to generate a dynamic interaction strategy including strategy priority marks and parameter adjustment amplitude values.
[0009] Optionally, optimizing the initial interaction strategy set through a strategy conflict resolution model to generate a dynamic interaction strategy including strategy priority marks and parameter adjustment amplitude values includes: Construct strategy conflict detection rules, detect the logical contradictions and execution feasibility conflicts among the dialogue guidance strategy, the task assistance strategy, and the emotional feedback strategy in the initial interaction strategy set, and generate corresponding conflict types; According to the detected conflict types, retrieve corresponding optimization strategies from a preset strategy optimization solution library to perform iterative optimization processing on the initial interaction strategy set until all conflicts are eliminated; Sort and integrate the optimized initial interaction strategy set according to the strategy priority marks and parameter adjustment amplitude values to generate the dynamic interaction strategy.
[0010] Optionally, displaying the multi-modal interaction content to the user terminal includes: Deploy a pre-generated cache pool of the multi-modal interaction content to an edge computing node connected to the user terminal, and render potential interaction content in this edge computing node; Obtain the current network latency measurement value of the user terminal and determine the corresponding data transmission mode; After adjusting the multi-modal interaction content according to the data transmission mode, send the adjusted multi-modal interaction content to the user terminal for display.
[0011] Optionally, the AI character interaction method further includes: Obtain the current game mode, where the game mode includes an online mode and a single-player mode; When the game mode is in the online mode, establish a network connection relationship between multiple user terminals, and render the multi-modal interaction content displayed on the current user terminal from the first perspective of the current user terminal, and synchronize the game status data to all user terminals participating in the online game in real time; When the game mode is in the single-player mode, disconnect the network connection with other user terminals, run the game logic on the local server of the user terminal, and store the corresponding game progress in the local storage medium.
[0012] In addition, to achieve the above object, the present invention also provides an interaction device, where the interaction device includes: a memory, a processor, and an AI character interaction program in the game platform stored on the memory and executable on the processor, and the AI character interaction program in the game platform is configured to implement the AI character interaction method in the game platform as described above.
[0013] In addition, to achieve the above object, the present invention also provides an interaction device including the interaction device as described above.
[0014] In an embodiment of the present invention, by obtaining the current game scene, game task information, and game character information, a first set of behavior parameters of the AI character is generated, and then the current operation behavior data and voice interaction data of the user are collected, and the operation behavior data and voice interaction data are subjected to intention recognition processing to generate a user behavior feature vector, where the user behavior feature vector includes the user's core intention, operation style feature, and emotional tendency. After the user behavior feature vector is generated, determine the relationship link between the user behavior feature vector and at least one of the game scene, task information, and character information, generate a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment amplitude, and then adjust the first set of behavior parameters according to the dynamic interaction strategy to generate a second set of behavior parameters, where the first set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient, and the second set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient, and then render and generate multi-modal interaction content matching the second set of behavior parameters, where the multi-modal interaction content includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient, and finally display the multi-modal interaction content on the user terminal to improve the smoothness and interaction experience of user interaction in the game platform. This embodiment can generate a dynamic interaction strategy and adjust the behavior parameters of the AI character according to the strategy, so as to render multi-modal interaction content matching the user behavior characteristics, making the interaction of the AI character more real and natural, and significantly improving the user's game immersion and satisfaction. Description of the Drawings
[0015] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0017] Figure 1 It is a schematic flowchart of the AI character interaction method within the game platform according to an embodiment of the present invention; Figure 2 It is a schematic flowchart of the AI character interaction method within the game platform according to another embodiment of the present invention; Figure 3 It is a schematic flowchart of the AI character interaction method within the game platform according to still another embodiment of the present invention; Figure 4 It is a schematic flowchart of the AI character interaction method within the game platform according to yet another embodiment of the present invention; Figure 5 It is a schematic flowchart of the AI character interaction method within the game platform according to still another embodiment of the present invention; Figure 6 It is a schematic flowchart of the AI character interaction method within the game platform according to another embodiment of the present invention; Figure 7 It is a schematic flowchart of the AI character interaction method within the game platform according to still another embodiment of the present invention; Figure 8 It is a schematic flowchart of the AI character interaction method within the game platform according to yet another embodiment of the present invention.
[0018] The realization of the objectives, functional features, and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Well-known modules, units, and their connections, links, communications, or operations are not shown or not described in detail. Moreover, the described features, architectures, or functions can be combined in any way in one or more embodiments. Those skilled in the art should understand that the following various embodiments are only for illustration, rather than for limiting the protection scope of the present invention. It can also be easily understood that the modules, units, or processing methods in the various embodiments described herein and shown in the drawings can be combined and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0020] For the definitions of various nouns or methods referred to in the following embodiments, except in cases where it is logically impossible to hold, the nouns or methods generally refer to the broad concepts that can be implemented on the premise of the content disclosed in the embodiments. Under such an understanding, all specific lower-level specific definitions of the nouns or methods should be regarded as the content of the present invention, and should not be narrowly understood or prejudicially interpreted on the grounds that the specific definition is not disclosed in the specification. Similarly, on the premise that it can be logically realized, the order of the steps in the method is flexible and changeable, and all specific lower-level specific definitions in the broad concepts of various nouns or methods belong to the scope of protection of the present invention.
[0021] Due to the unsmooth user interaction process and poor interaction experience in the existing game platforms. There are obvious limitations in the AI character interaction methods in traditional games: First, the understanding of the user's intention by the AI character is only based on simple keyword matching, and it is unable to accurately capture the core intention, operation style, and emotional tendency contained in the user's operation behavior and voice interaction; Second, the interaction strategy adopts a fixed template and lacks the ability to dynamically adjust according to the game scene, task progress, and character attributes; Third, the form of presenting the interaction content is single, and it is difficult to achieve the spatio-temporal synchronization of multi-modal elements such as voice, actions, and special effects. These problems lead to the rigid and stereotyped behavior performance of the AI character during the game process, and the interaction with the user lacks a sense of reality and immersion, seriously affecting the game experience. Especially when the network environment fluctuates, the existing technology cannot ensure the smooth transmission and presentation of multi-modal interaction content, greatly exacerbating the sense of fragmentation of the interaction experience.
[0022] The main solution of the embodiment of the present application is as follows: By obtaining the current game scene, game task information, and game character information, a first set of behavior parameters for the AI character is generated. Then, the current operation behavior data and voice interaction data of the user are collected, and the operation behavior data and voice interaction data are subjected to intention recognition processing to generate a user behavior feature vector, where the user behavior feature vector includes the user's core intention, operation style features, and emotional tendency. After the user behavior feature vector is generated, the relationship link between the user behavior feature vector and at least one of the game scene, task information, and character information is determined, and a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment range is generated. Then, the first set of behavior parameters is adjusted according to the dynamic interaction strategy to generate a second set of behavior parameters. The first set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient, and the second set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient. Then, multimodal interaction content matching the second set of behavior parameters is rendered and generated, where the multimodal interaction content includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient. Finally, the multimodal interaction content is displayed on the user terminal.
[0023] In this embodiment, for the convenience of description, the following will be described with an interaction device as the execution subject.
[0024] The present application provides a solution to improve the smoothness and interaction experience of user interactions within a game platform. This embodiment can generate a dynamic interaction strategy and adjust the behavior parameters of the AI character according to the strategy, thereby rendering multimodal interaction content that matches the user's behavior characteristics, making the interaction of the AI character more real and natural, and significantly improving the user's game immersion and satisfaction.
[0025] For this reason, the present invention proposes an AI character interaction method within a game platform. It can be understood that an interaction device for storing and executing the following method is provided within the interaction device, and the interaction device can be implemented using a main controller, such as an MCU (Microcontroller Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), an SOC (System On Chip), etc.
[0026] In the prior art, there are generally problems of high operation latency and low intention recognition accuracy in the user interaction process of game platforms. Traditional systems rely on preset scripts to drive the behavior of AI characters and cannot adjust interaction strategies according to real-time game scenarios and user dynamics. In typical scenarios, when users complete multiplayer dungeon tasks, there is a timing misalignment between voice commands and interface operations, resulting in the task guidance provided by AI characters not matching the current combat stage.
[0027] To solve the above problems, it is necessary to break through the static strategy framework and build a dynamic adaptive multi-modal interaction mechanism. It is found in the design process that it is difficult to capture users' composite intentions through single data source analysis, and it is necessary to fuse voice, operation behavior, and scene context for joint modeling.
[0028] Based on the above content, referring to Figure 1 , in an embodiment of the present invention, the AI character interaction method in the game platform includes steps S100 - S700, where: S100. Obtain the current game scene, game task information, and game character information, and generate a first set of behavior parameters for the AI character; S200. Collect the current operation behavior data and voice interaction data of the user; S300. Perform intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector; S400. Determine the relationship link between the user behavior feature vector and at least one of the game scene, the task information, and the character information to generate a dynamic interaction strategy with interaction strategy priority and parameter adjustment amplitude; S500. Adjust the first set of behavior parameters according to the dynamic interaction strategy to generate a second set of behavior parameters, where the first set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient; S600. Render and generate multi-modal interaction content matching the second set of behavior parameters, where the multi-modal interaction content includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient; S700. Display the multi-modal interaction content to the user terminal.
[0029] Among them, the user behavior feature vector includes the user's core intention, operation style features, and emotional tendency. The user behavior feature vector refers to the set of user interaction features extracted through multi-dimensional data analysis, which can be achieved by acoustic feature normalization processing and operation action semantic segmentation technology, and is used to accurately represent the user's core intention and emotional state. The dynamic interaction strategy refers to the interaction plan generated based on the real-time context, which can be specifically achieved by constructing a dynamic relationship graph to calculate the multi-dimensional association strength, and is used to guide the dynamic adjustment of the AI character's behavior parameters. The multi-modal interaction content refers to the composite output that integrates voice, actions, and visual effects, and can specifically adopt a multi-modal synchronization engine to achieve timestamp alignment to ensure the spatio-temporal consistency of the interaction content in different sensory channels.
[0030] Among them, the interaction device obtains the scene terrain data, task objective list, and NPC attribute parameters through the game engine interface, and initializes the basic behavior mode of the AI character. For example, when the user executes a shooting game rescue mission, the voice command "Cover the left side" and the operation data of quickly switching weapons are synchronously collected. By analyzing the hot zone distribution, it is identified that the user frequently clicks on the left side cover area, and combined with voice recognition to extract the tactical instruction keywords, a feature vector containing the tactical deployment intention and aggressive operation style is generated. The interaction device performs correlation analysis on this feature vector with the task progress stage and the enemy firepower distribution, and dynamically increases the left-side co-defense frequency and firepower suppression intensity of the AI teammate. The adjusted behavior parameters drive the generation of a voice prompt "Blocking the left passage", an animation of the character quickly moving to the cover position, and a highlighted boundary special effect in the left area, forming a coordinated multi-modal feedback.
[0031] Compared with the prior art, the traditional method uses a fixed priority strategy to process user input, and it is easy to generate strategy conflicts when the voice command and the operation behavior point to different targets. In this embodiment, through the spatio-temporal alignment model, the fusion and analysis of multi-source data are realized, and the correlation degree between the user intention and the game elements is quantified by combining the dynamic relationship graph, effectively solving the problem of collaborative processing of cross-modal instructions. In the open-world exploration scenario, when the user asks "Where is the treasure" while operating the character to move towards the map marker, the interaction device can accurately identify the redundant guidance requirement and automatically reduce the task prompt frequency.
[0032] Through the above technical means, this embodiment realizes the context awareness and adaptive interaction of the AI character in the game environment. In real-time strategy games, when the user simultaneously conducts voice command and multi-unit micro-operation, the interaction device can accurately identify the core tactical intention and dynamically adjust the autonomous behavior parameters of the AI unit. In role-playing tasks, according to the emotional tendency of the user's dialogue, the facial expressions and voice intonations of the NPC are automatically matched, enhancing the immersion and naturalness of the interaction process.
[0033] In this embodiment, by obtaining the current game scene, game task information, and game character information, a first set of behavior parameters for the AI character is generated. Then, the current operation behavior data and voice interaction data of the user are collected, and the operation behavior data and voice interaction data are subjected to intent recognition processing to generate a user behavior feature vector, where the user behavior feature vector includes the user's core intent, operation style features, and emotional tendency. After the user behavior feature vector is generated, the relationship link between the user behavior feature vector and at least one of the game scene, task information, and character information is determined, and a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment range is generated. Then, the first set of behavior parameters is adjusted according to the dynamic interaction strategy to generate a second set of behavior parameters, where the first set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient, and the second set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient. Then, multi-modal interaction content matching the second set of behavior parameters is rendered and generated, where the multi-modal interaction content includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient. Finally, the multi-modal interaction content is displayed to the user terminal to improve the smoothness and interaction experience of the user interaction within the game platform. This embodiment can generate a dynamic interaction strategy and adjust the behavior parameters of the AI character according to the strategy, thereby rendering multi-modal interaction content matching the user's behavior characteristics, making the interaction of the AI character more real and natural, and significantly improving the user's game immersion and satisfaction.
[0034] Optionally, referring to Figure 2 , another embodiment of the present invention provides an AI character interaction method within a game platform. Based on the above Figure 1 shown embodiment, the processing of performing intent recognition on the operation behavior data and voice interaction data to generate a user behavior feature vector includes steps S310 - S340, where: S310. After performing acoustic feature standardization processing on the voice interaction data, input the processed voice interaction data into a speech recognition engine to convert the processed voice interaction data into a text statement sequence with timestamp marks; S320. Perform operation action semantic segmentation on the operation behavior data based on the hot zone distribution, frequency - intensity grading, and validity marking of the task progress status of the display interface of the user terminal to generate a corresponding operation action statement sequence; S330. Construct a spatio - temporal alignment model for the text statement sequence and the operation action statement sequence, perform timestamp dynamic matching, game object association mapping, and logical consistency detection to generate an intermediate feature set including semantic keyword weight distribution, decision - making preference intensity index, and emotional tendency index; S340. Infer the intention of the intermediate feature set based on the context state matrix with dynamic update, determine the intention probability distribution, and generate a user behavior feature vector corresponding to the intention probability distribution.
[0035] Among them, acoustic feature normalization refers to the processing of noise reduction, intonation unification, and dialect adaptation for speech data. It can also be implemented using noise reduction algorithms based on deep neural networks, Mel-frequency cepstral coefficient analysis, and dialect speech database matching algorithms, which are used to eliminate environmental interference and improve the recognition compatibility of different accents. Operational action semantic segmentation refers to the structural analysis of operational behaviors based on the hot zone distribution of user operations, operation frequency and intensity, and task progress effectiveness. It can be achieved through an interface hot zone tracking module, pressure sensor data acquisition, and task status marking interaction device, which are used to convert discrete operational behaviors into sequence data with semantic information. The spatio-temporal alignment model refers to establishing the correlation between speech text and operational actions in the time and space dimensions. It can be implemented using dynamic time warping algorithms, game object recognition engines, and logic rule validators, which are used to ensure the consistency of speech commands and operational behaviors in terms of timing and logic. The context state matrix refers to a dynamic data set that integrates the current task objective, user historical behavior patterns, and scene object attributes. It can be achieved through a task status tracker, user behavior log analyzer, and scene attribute database, which are used to provide real-time context support for intention inference.
[0036] Among them, the speech data is converted into a timestamped text sequence after noise reduction, intonation normalization, and dialect adaptation, eliminating the impact of environmental noise and accent differences on the recognition accuracy. The operational behavior data determines the user's attention area according to the interface hot zone distribution, combines the operation frequency and intensity to grade and quantify the operation intensity, and filters out invalid operations through the task progress effectiveness marking, and is segmented into semantic action statements. The text sequence and action statements are dynamically matched and aligned in time sequence through timestamps, and a mapping relationship between speech commands and operation targets is established using a game object recognition engine. Contradictory data is excluded through a logic rule validator, forming intermediate features including keyword weights, decision preferences, and emotional tendencies. The context matrix with dynamic update integrates the current task phase objective, user historical behavior statistical features, and scene object attribute relationships, calculates the intention distribution through a probability model, and finally generates a behavior feature vector reflecting the user's real-time intention.
[0037] Compared with the prior art, traditional methods usually process speech and operation data separately, resulting in isolated intent recognition and lack of context association. Speech recognition in the prior art does not consider dialect differences and environmental noise interference. Operation behavior parsing does not combine the effectiveness of task progress and the distribution of interface hotspots. Temporal and spatial alignment only uses simple timestamp matching without establishing a mapping of game object associations. In this embodiment, multi-source data preprocessing is achieved through acoustic standardization and operation semantic segmentation, a spatio-temporal alignment model is constructed to integrate temporal and logical relationships, and intent inference is realized by combining a dynamic context matrix, solving the problem of intent deviation caused by isolated data processing.
[0038] Through the above technical means, this embodiment can effectively integrate multi-dimensional data of user speech and operation behavior, eliminate the influence of environmental interference and operation noise on intent recognition, and accurately extract the weights of semantic keywords and decision preference features. Through intent probability inference supported by dynamic context, the generated user behavior feature vector can accurately reflect the core intent and behavior pattern of the user in the current game scenario, providing a reliable data basis for the generation of subsequent interaction strategies, and significantly improving the response accuracy and scene adaptability of the game AI character to user intents.
[0039] Optionally, referring to Figure 3 , another embodiment of the present invention provides an AI character interaction method in a game platform. Based on the above Figure 1 shown embodiment, rendering and generating multi-modal interaction content matching the second set of behavior parameters includes steps S610 - S640, where: S610. Based on the dialogue response mode in the second set of behavior parameters, call a speech synthesis engine to generate speech waveform data with emotional prosody features; S620. According to the task guidance intensity value in the second set of behavior parameters, extract basic action primitives from a preset action library, and through a physical simulation engine, mix and optimize the basic action primitives to generate three-dimensional character action data matching the task information, and generate an interface highlight special effect positively correlated with the guidance intensity value; S630. Based on the emotional expression coefficient in the second set of behavior parameters, generate a facial micro-expression fusion weight, and control a particle special effect engine to generate an environmental feedback special effect matching the emotional dimension based on the facial micro-expression fusion weight; S640. Input the speech waveform data, the three-dimensional character action data, the interface highlight special effect, and the environmental feedback special effect into a multi-modal synchronization engine to generate multi-modal interaction content with spatio-temporal feature alignment.
[0040] Among them, the dialogue response mode refers to the speech output logic matched according to the user's intention, which can be implemented by a speech synthesis model based on sentiment classification. Its function is to ensure the semantic consistency between the speech content and the user's current intention. The task guidance intensity value refers to the quantified task promotion strength parameter, which can be calculated by using a sliding window to statistically analyze the relationship between the task completion degree and the operation frequency. Its function is to dynamically adjust the amplitude of the character's actions and the intensity of the interface prompts. The emotion expression coefficient refers to the weight distribution parameter of multi-dimensional emotion feedback, which can be generated by combining an emotion dimension mapping matrix with real-time emotion recognition results. Its function is to coordinate the synchronous expression of facial expressions and environmental special effects. The multi-modal synchronization engine refers to a spatio-temporal alignment data processing module, which can be implemented by combining a dynamic time warping algorithm with an event-driven mechanism. Its function is to eliminate the time offset and logical conflicts of different modal data.
[0041] Among them, during the process of generating multi-modal interaction content, the speech synthesis engine adjusts the fundamental frequency and speech rate parameters according to the sentiment classification results in the dialogue response mode to generate waveform data that matches the user's intention. The physical simulation engine optimizes the trajectory of the extracted basic action primitives through an inverse kinematics algorithm to generate a three-dimensional action sequence that conforms to the task guidance intensity. At the same time, it linearly adjusts the emission frequency of the interface elements according to the guidance intensity value. The facial micro-expression fusion weight superimposes multiple basic expressions according to the emotion coefficient ratio through a blend shape technique, and the particle special effect engine adjusts the particle movement trajectory and color gradient mode according to the emotion dimension parameters. The multi-modal synchronization engine captures the output signals of each module through an event-driven mechanism and uses a dynamic time warping algorithm to align the time axes of the speech waveform, character actions, and environmental special effects to ensure that all modal data reaches a synchronous output state at the interaction trigger moment.
[0042] Compared with the prior art, the traditional method simply superimposes the independently generated data of each modality, which easily leads to the disconnection between speech and actions, and the incoordination between special effects and expressions. In this embodiment, by establishing a parameter linkage mechanism and a multi-modal synchronization engine, spatio-temporal associations between various elements are established during the data generation stage, realizing the dynamic binding of action primitives and task guidance intensity, the dimension mapping of emotion coefficients and environmental special effects, and eliminating cross-modal response delays through a unified time reference.
[0043] Through the above technical means, this embodiment solves the problem of spatio-temporal asynchrony during the generation of multi-modal interaction content, ensures the logical consistency between speech guidance and character actions during the task promotion process, realizes the dynamic matching of facial expression changes and environmental special effects, and effectively improves the overall perceptual fluency of the user for the interaction content. The application of the multi-modal synchronization engine enables different interaction elements to form a synergistic effect in terms of trigger timing and performance intensity, avoiding the time misalignment phenomenon between visual cues and speech feedback in the traditional method, and enhancing the effectiveness of task guidance and the immersion of emotion expression.
[0044] Optionally, referring to Figure 4 , another embodiment of the present invention provides an AI character interaction method in a game platform. Based on the above Figure 3 shown embodiment, the method of inputting the voice waveform data, the three-dimensional character action data, the interface highlighting special effects, and the environmental feedback special effects into a multimodal synchronization engine to generate multimodal interaction content with spatio-temporal feature alignment includes steps S641 - S645, where: S641. Add a first-level timestamp marker to the voice waveform data according to the timing requirements of the dialogue response mode; S642. Add a second-level timestamp marker to the three-dimensional character action data based on the physical simulation logic of the task guidance intensity; S643. Add a third-level timestamp marker to the environmental feedback special effects according to the dynamic change curve of the emotion expression coefficient; S644. Align the first-level, second-level, and third-level timestamp markers through a dynamic offset compensation algorithm to generate a unified reference timeline; S645. Input the voice waveform data, the three-dimensional character action data, the interface highlighting special effects, and the environmental feedback special effects into the multimodal synchronization engine according to the unified reference timeline for spatio-temporal feature alignment processing.
[0045] Among them, the first-level timestamp marker refers to a synchronization marker set based on the voice output timing characteristics of the dialogue response mode, which can be implemented by combining voice segmentation processing technology with a voice activity detection algorithm, and is used to ensure that the natural rhythm of the voice waveform data matches the user operation in real time. The second-level timestamp marker refers to the time encoding of the character action sequence based on the physical simulation logic, which can be generated by combining a rigid body dynamics model with a motion trajectory interpolation algorithm, so that the three-dimensional character actions conform to the physical motion laws and are synchronized with the change of the task guidance intensity. The third-level timestamp marker refers to a special effect trigger marker generated according to the emotion expression curve, which can be implemented by combining a time series prediction model with an emotion dimension mapping rule, and ensures that the particle emission frequency of the environmental feedback special effects is consistent with the change trend of the emotion tendency. The dynamic offset compensation algorithm refers to a synchronization correction method for eliminating the time error of the multimodal data stream, which can be implemented by combining a clock drift estimation model with an adaptive delay compensation mechanism, and generates a unified synchronization timeline by real-time monitoring the time offset of each data stream and dynamically adjusting the time reference.
[0046] Among them, when the first-level timestamp marks are added to the speech waveform data, the speech segment boundaries are divided through speech endpoint detection technology, and the time mark intervals are set according to the rhythm requirements of the dialogue response mode. For example, an extended mark is set at the end of an interrogative sentence. When the second-level timestamp marks are added to the 3D character action data, time encoding is performed on the key action frames according to the skeletal movement trajectory calculated by the physics engine. For example, when the character performs a jumping action, different time marks are set for the takeoff, mid-air, and landing stages. When the third-level timestamp marks are added to the environmental feedback special effects, according to the numerical change rate of the emotion expression coefficient, special effect intensity adjustment marks are set at the inflection points of the curve. For example, a particle burst mark is set when the emotion coefficient suddenly increases. The dynamic offset compensation algorithm calculates the time deviation amounts of each stream by monitoring the transmission delays of the speech, action, and special effect data streams in real time. For example, when the speech stream is 5 milliseconds faster than the action stream, the time marks of the action stream are moved forward for compensation to generate a unified time reference axis. The multimodal synchronization engine performs frame synchronization processing on each data stream according to this time axis. For example, at a specific time point, the speech waveform playback, character action execution, and interface special effect rendering are triggered simultaneously.
[0047] Compared with the prior art, traditional multimodal synchronization methods usually adopt a single timestamp marking system and cannot adapt to the differences in the generation mechanisms of different data streams. For example, fixed frame rate marking is used for speech data while action data depends on physical simulation calculations, which easily leads to out-of-sync lip movements and speech. In the prior art, environmental special effects mostly adopt preset trigger conditions and cannot dynamically adapt to changes in emotion expression. For example, special effect delays may occur when the player's emotions are intense. In this embodiment, by establishing a hierarchical timestamp system, different marking rules are set according to the generation principles of different modal data, and a dynamic compensation mechanism is combined to eliminate cross-modal time errors. For example, when the action data lags due to physical simulation calculation delays, the presentation timing of other data streams is automatically adjusted.
[0048] Through the above technical means, this embodiment solves the problem of spatio-temporal out-of-sync of multimodal data caused by differences in generation mechanisms, making the speech output rhythm accurately match the character lip movement changes, the task guiding actions coordinated with the interface prompt special effects, and the intensity of the environmental particle special effects adjusted in real time according to emotions. It avoids technical defects such as the character starting to move only after the speech playback is completed and the emotional special effects disappearing in advance, affecting the immersion in the traditional methods, and realizes the precise synchronous presentation of multi-dimensional interaction elements.
[0049] Optionally, referring to Figure 5 , another embodiment of the present invention provides an AI character interaction method in a game platform, based on the above Figure 1In the illustrated embodiment, determining the relationship link between the user behavior feature vector and at least one of the game scenario, the task information, and the character information to generate a dynamic interaction strategy with interaction strategy priorities and parameter adjustment amplitudes includes steps S410 - S440, where: S410. Construct a dynamic relationship graph including game scenario objects, task target nodes, and character attribute nodes; S420. Calculate the multi - dimensional association strength between the user behavior feature vector and at least one type of node in the dynamic relationship graph; S430. Generate an initial interaction strategy set based on the multi - dimensional association strength; S440. Optimize the initial interaction strategy set through a strategy conflict resolution model to generate a dynamic interaction strategy including strategy priority marks and parameter adjustment amplitude values.
[0050] Among them, the dynamic relationship graph refers to a data structure that establishes associations between interactive objects, task targets, and character attributes in a game scenario in the form of nodes. It can be implemented using graph database or knowledge graph technology and is used to integrate multi - dimensional information in the game environment in real - time. The multi - dimensional association strength refers to a similarity index calculated between the user behavior feature and the graph nodes through a vector space model. It can be implemented using cosine similarity algorithm or neural network embedding technology and is used to quantify the association degree between the user intention and game elements. The initial interaction strategy set refers to a candidate strategy combination generated based on the association strength. It can be implemented using a rule engine or probability model and is used to cover interaction requirements in different scenarios. The strategy conflict resolution model refers to an optimization module that detects logical contradictions between strategies. It can be implemented using constraint satisfaction algorithms or reinforcement learning models and is used to eliminate conflicts during strategy execution.
[0051] Among them, by constructing a dynamic relationship graph to integrate game scenario, task target, and character attribute information, it provides a structured data basis for subsequent association strength calculation. When calculating the multi - dimensional association strength between the user behavior feature vector and the graph nodes, a vector similarity algorithm is used to quantify the association degree between the user intention and game elements, generating an initial strategy set covering different interaction dimensions. Logical contradictions between dialogue guidance, task assistance, and emotional feedback strategies are identified through conflict detection rules, and the conflict resolution solutions in the optimized strategy library are called for iterative processing. Finally, the optimized strategies are sorted and integrated according to priorities and adjustment amplitudes to form a conflict - free dynamic interaction strategy.
[0052] Compared with the prior art, traditional methods usually generate strategies based on single-dimensional user behavior data, lacking dynamic correlation analysis of game scenarios, tasks, and character attributes, resulting in execution conflicts easily occurring when multiple strategies are executed in parallel. In this embodiment, multi-dimensional information integration is achieved by constructing a dynamic relationship graph, and the strategies are optimized through an interactive device by combining a conflict resolution model, which can effectively coordinate the execution logic between different strategies.
[0053] Through the above technical means, this embodiment solves the logical conflict problem generated during the parallel execution of multiple strategies, avoids interaction interruption or behavior chaos caused by strategy contradictions, and improves the fluency of the interaction process and the consistency of strategy execution.
[0054] Optionally, referring to Figure 6 , another embodiment of the present invention provides an AI character interaction method in a game platform. Based on the above Figure 5 shown embodiment, the initial interaction strategy set is optimized through a strategy conflict resolution model to generate a dynamic interaction strategy including a strategy priority mark and a parameter adjustment amplitude value, including steps S441 - S443, where: S441. Construct a strategy conflict detection rule to detect the logical contradiction and execution feasibility conflict among the dialogue guidance strategy, task assistance strategy, and emotional feedback strategy in the initial interaction strategy set, and generate corresponding conflict types; S442. According to the detected conflict types, retrieve corresponding optimization strategies from a preset strategy optimization solution library to perform iterative optimization processing on the initial interaction strategy set until all conflicts are eliminated; S443. Sort and integrate the optimized initial interaction strategy set according to the strategy priority mark and the parameter adjustment amplitude value to generate the dynamic interaction strategy.
[0055] Among them, the strategy conflict detection rule refers to a verification mechanism for identifying logical contradictions and execution feasibility conflicts between different strategies, which can be implemented by semantic logic verification, resource occupancy simulation, or execution timing analysis to detect implicit conflicts between strategies. The preset strategy optimization solution library refers to a database storing the mapping relationship between conflict types and optimization strategies, which can be implemented by a hash index structure based on conflict labels or a graph neural network matching algorithm to quickly locate applicable optimization strategies. The strategy priority mark refers to metadata for identifying the execution order of strategies, which can be implemented by a weight scoring mechanism or a dependency relationship graph to determine the priority order during the parallel execution of multiple strategies. The parameter adjustment amplitude value refers to a quantitative index of the degree of modification of strategy parameters, which can be implemented by a gradient descent algorithm or a fuzzy logic controller to control the boundary range of strategy adjustment.
[0056] Among them, by constructing policy conflict detection rules, logical contradiction detection is carried out on the dialogue guidance policy, task assistance policy, and emotional feedback policy in the initial interactive policy set. For example, when the task assistance policy requires forced progress of the task, and the emotional feedback policy detects negative emotions of the user, a feasibility conflict is triggered. After detecting the conflict type, an optimization policy corresponding to the conflict type is called from the preset policy optimization solution library. For example, for the conflict between task promotion and emotion appeasement, an optimization policy combining task goal splitting and emotional compensation mechanism is adopted. Through iterative optimization processing, such as performing conflict detection again after the first optimization, to ensure that all conflicts are eliminated. The optimized policy set is given a priority mark. For example, the task assistance policy is set to a high priority, and the emotional feedback policy is set to dynamically adjust the priority. At the same time, the range of policy parameter modification is restricted according to the parameter adjustment amplitude value, and finally a conflict-free dynamic interactive policy is generated.
[0057] Compared with the prior art, traditional methods adopt single policy optimization or manual rule setting of priorities, and cannot interactively solve the implicit conflicts during the collaborative execution of multiple policies. For example, only by adjusting weights to cover up contradictions, resulting in logical breaks in the interaction process. The prior art lacks an accurate matching mechanism for optimization policies for conflict types, and often uses fixed priority sorting, resulting in rigid policies. This embodiment realizes the root cause elimination of policy conflicts through a hierarchical conflict detection and iterative optimization mechanism, while retaining the integrity and execution flexibility of the policy system.
[0058] Through the above technical means, this embodiment solves the problem of interaction process jamming caused by logical contradictions during the collaborative execution of multiple policies. By identifying conflict types and calling targeted optimization policies, it avoids parameter out-of-control during the policy adjustment process, ensures that the dynamically generated interactive policies maintain the balance of task promotion efficiency and emotional feedback while eliminating internal conflicts, and improves the interaction coherence of users in complex game scenarios.
[0059] Optionally, referring to Figure 7 , another embodiment of the present invention provides an AI character interaction method in a game platform. Based on the above Figure 1 shown embodiment, the displaying the multimodal interaction content to the user terminal includes steps S710 - S730, where: S710. Deploy the pre-generated cache pool of the multimodal interaction content to an edge computing node connected to the user terminal, and render potential interaction content in this edge computing node; S720. Obtain the current network delay measurement value of the user terminal and determine the corresponding data transmission mode; S730. After adjusting the multimodal interaction content according to the data transmission mode, send the adjusted multimodal interaction content to the user terminal for display.
[0060] Establish a dynamic bitrate adjustment mechanism, and adaptively select the following transmission modes according to the measured network latency: when the latency is less than 50 milliseconds, transmit the complete multi-modal data stream; when the latency is between 50 milliseconds and 150 milliseconds, enable the progressive loading strategy; when the latency is greater than or equal to 150 milliseconds, switch to the lightweight interaction protocol and only retain the key semantic data.
[0061] Among them, the pre-generated cache pool refers to storing the rendered interactive content data blocks in the edge computing nodes in advance, which can be realized by combining distributed storage technology with the rendering task queue management, and is used to reduce the computational pressure and data transmission volume of real-time rendering. The edge computing node refers to a server cluster deployed in the vicinity of the user terminal, which can be realized through the content delivery network architecture, and is used to shorten the data transmission path and reduce the network latency. The measured network latency value refers to the communication response time index between the user terminal and the server, which can be realized by combining the round-trip time measurement algorithm with the packet loss rate detection, and is used to evaluate the current network transmission quality. The data transmission mode refers to a combination of data transmission strategies dynamically adjusted according to the network state, which can be realized by combining protocol stack parameter configuration with data compression algorithms, and is used to adapt to the transmission requirements under different network conditions. The dynamic bitrate adjustment mechanism refers to a decision-making model that automatically switches the transmission strategy based on the network latency amount, which can be realized by combining multi-level threshold triggering rules with priority queue scheduling, and is used to balance data integrity and transmission efficiency.
[0062] Among them, the pre-generated cache pool generates potential interactive content in advance through the local rendering ability of the edge computing node, avoiding the high latency problem caused by cloud remote rendering. The measured network latency value is monitored in real time and classified into different levels, triggering the corresponding transmission mode: in the low latency scenario, transmit the complete data stream to ensure the details of the interactive content; in the medium latency scenario, give priority to transmitting the core interactive elements and gradually load the auxiliary content; in the high latency scenario, only retain the key semantic data to maintain the basic interaction logic. The dynamic bitrate adjustment mechanism divides the network state interval through multi-level thresholds, ensures that the transmission strategy matches the network fluctuations in real time, and at the same time reduces the dependence on cloud data by combining the local cache of the edge node.
[0063] Compared with the prior art, the traditional method adopts a fixed bitrate transmission mode, which cannot cope with the latency changes caused by network fluctuations, and is prone to data congestion or content loss. In the prior art, the rendering of interactive content usually depends on the cloud server, resulting in a too long transmission path and difficulty in real-time response. In this embodiment, by deploying edge computing nodes to achieve local rendering and caching, combined with the dynamic bitrate adjustment mechanism, the data transmission path is effectively shortened and the network state changes are adapted.
[0064] Through the above technical means, this embodiment solves the problems of transmission delay or lag of interactive content caused by network latency, ensuring the coherence and real-time nature of multimodal interactive content under different network conditions. Through edge node pre-rendering and dynamic transmission strategies, the data transmission load between the cloud and the terminal is reduced, while the risk of network congestion in the fixed bitrate mode is avoided. When the network state fluctuates, the amount of data transmitted is adaptively adjusted, maintaining the availability of the core interactive functions while optimizing the user experience in high-bandwidth scenarios.
[0065] Optionally, referring to Figure 8 , another embodiment of the present invention provides an AI character interaction method within a game platform. Based on the embodiment shown in Figure 1 above, the AI character interaction method further includes steps S800 - S1000, where: S800. Obtain the current game mode, where the game mode includes an online mode and a single-player mode; S900. When the game mode is in the online mode, establish a network connection relationship between multiple user terminals, and render the multimodal interactive content displayed on the current user terminal from the first perspective of the current user terminal, and synchronize the game state data to all user terminals participating in the online game in real time; S1000. When the game mode is in the single-player mode, disconnect the network connection with other user terminals, run the game logic on the local server of the user terminal, and store the corresponding game progress in the local storage medium.
[0066] Among them, the game mode refers to the state classification of the operating environment, which can be implemented by the user actively selecting or the interactive device automatically detecting, and is used to trigger different network connection strategies and data processing mechanisms. The network connection relationship refers to the communication link topology structure between terminals, which can be established using the P2P protocol or the central server architecture, and is used to ensure the data interaction requirements among multiple users in the online mode. The first-perspective rendering refers to the display method centered on the current user operation interface, which can be realized through a view-locking algorithm and is used to ensure the primary and secondary hierarchical relationship of the interactive content. The local server refers to the logical operation unit deployed in the user terminal, which can be implemented using lightweight virtualization technology and is used to independently process the game operation and storage requirements in the single-player mode.
[0067] Among them, when the interaction device detects that the user enters the online mode, it automatically triggers the network connection module to build a communication channel among multiple terminals. For example, after confirming the online status of each terminal through the heartbeat packet detection mechanism, the distributed synchronization protocol is used to render the multi-modal interaction content of the current user with priority, and at the same time, the game state data is compressed and encapsulated and then transmitted to other participating terminals. During this process, the perspective locking algorithm continuously tracks the user operation interface to ensure that the three-dimensional scene rendering and interaction feedback are presented from the first perspective. When switching to the single-player mode, the network connection module actively releases the communication resources, and the local server takes over the game logic operation. For example, the game process running state is maintained through the memory resident technology, and at the same time, the key progress data is encrypted and written into the local storage medium.
[0068] Compared with the prior art, the traditional solution requires manual adjustment of the network configuration when switching modes and cannot maintain data consistency. However, in this embodiment, through the dynamic network topology reconstruction and localization processing mechanism, real-time synchronization of multi-terminal data in the online mode and offline autonomous operation in the single-player mode are achieved. In the prior art, the online mode relies on a fixed server architecture, resulting in high latency. In this embodiment, the distributed communication protocol is adopted to reduce the data transmission latency; in the single-player mode, the traditional solution relies on external storage, which has security risks. In this embodiment, the data security is enhanced through local encrypted storage.
[0069] Through the above technical means, this embodiment solves the problem of interaction chaos caused by asynchronous multi-user perspectives in the online mode, avoids resource waste caused by redundant network connections in the single-player mode, and at the same time ensures the integrity and recoverability of the game progress through the local storage mechanism. The real-time synchronization mechanism in the online mode ensures the state consistency of all terminals, and the independent operation logic in the single-player mode improves the operation response speed. The seamless switching between the two modes optimizes the overall interaction experience.
[0070] The present invention also proposes an interaction device, which includes: a memory, a processor, and an AI character interaction program in the game platform stored on the memory and executable on the processor. The AI character interaction program in the game platform is configured to implement the AI character interaction method in the game platform as described above.
[0071] It should be noted that since the interaction device of the present invention is based on the above AI character interaction method in the game platform, therefore, the embodiments of the interaction device of the present invention include all the technical solutions of all the embodiments of the above AI character interaction method in the game platform, and the achieved technical effects are also exactly the same, which will not be elaborated here.
[0072] The present invention also proposes an interaction device, which includes the interaction device as described in the above embodiment.
[0073] It should be noted that since the interaction device of the present invention is based on the above-mentioned interaction device, the embodiments of the interaction device of the present invention include all the technical solutions of all the embodiments of the above-mentioned interaction device, and the achieved technical effects are exactly the same, so they will not be elaborated here.
[0074] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0075] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.
[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0077] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. An AI character interaction method within a game platform, characterized in that The described AI character interaction method includes: Obtaining the current game scene, game task information, and game character information, and generating a first set of behavior parameters for the AI character; Collecting the current operation behavior data and voice interaction data of the user; Performing intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, where the user behavior feature vector includes the user's core intention, operation style features, and emotional tendency; Determining the relationship link between the user behavior feature vector and at least one of the game scene, the task information, and the character information to generate a dynamic interaction strategy with interaction strategy priorities and parameter adjustment amplitudes; Adjusting the first set of behavior parameters according to the dynamic interaction strategy to generate a second set of behavior parameters, where the first set of behavior parameters includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient; Rendering and generating multimodal interaction content that matches the second set of behavior parameters, where the multimodal interaction content includes a dialogue response mode, a task guidance intensity, and an emotional expression coefficient; Displaying the multimodal interaction content to the user terminal.
2. The AI character interaction method within the game platform according to claim 1, wherein The performing intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector includes: Performing acoustic feature normalization processing on the voice interaction data and inputting the processed voice interaction data into a speech recognition engine to convert the processed voice interaction data into a text statement sequence with timestamp marks; Performing operation action semantic segmentation on the operation behavior data based on the hot zone distribution, frequency intensity grading, and validity marking of the task progress status of the display interface of the user terminal to generate a corresponding operation action statement sequence; Constructing a spatio-temporal alignment model for the text statement sequence and the operation action statement sequence, performing timestamp dynamic matching, game object association mapping, and logical consistency detection, and generating an intermediate feature set including semantic keyword weight distribution, decision preference intensity index, and emotional tendency index; Performing intention inference on the intermediate feature set based on a dynamically updated context state matrix, determining an intention probability distribution, and generating a user behavior feature vector corresponding to the intention probability distribution, where the context state matrix includes the current task stage target, the user's historical behavior statistical features, and the scene object attribute relationship graph.
3. The AI character interaction method in the game platform according to claim 1, characterized in that, The rendering and generating multimodal interaction content that matches the second set of behavior parameters includes: Based on the dialogue response mode in the second set of behavior parameters, calling a speech synthesis engine to generate voice waveform data with emotional prosody features; According to the task guidance intensity value in the second set of behavior parameters, extracting basic action primitives from a preset action library, and performing action sequence hybrid optimization on the basic action primitives through a physical simulation engine to generate three-dimensional character action data that matches the task information, and generating an interface highlight special effect that is positively correlated with the guidance intensity value; Based on the emotional expression coefficient in the second set of behavior parameters, generating a facial micro-expression fusion weight, and controlling a particle special effect engine to generate an environment feedback special effect that matches the emotional dimension based on the facial micro-expression fusion weight; Input the voice waveform data, the 3D character action data, the interface highlighting special effects, and the environmental feedback special effects into a multimodal synchronization engine to generate multimodal interaction content with aligned spatio-temporal features.
4. The AI character interaction method within the game platform according to claim 3, wherein, The step of inputting the voice waveform data, the 3D character action data, the interface highlighting special effects, and the environmental feedback special effects into a multimodal synchronization engine to generate multimodal interaction content with aligned spatio-temporal features includes: Add first-level timestamp marks to the voice waveform data according to the timing requirements of the dialogue response mode; Add second-level timestamp marks to the 3D character action data based on the physical simulation logic of the task guidance intensity; Add third-level timestamp marks to the environmental feedback special effects according to the dynamic change curve of the emotion expression coefficient; Align the first-level, second-level, and third-level timestamp marks through a dynamic offset compensation algorithm to generate a unified reference timeline; Input the voice waveform data, the 3D character action data, the interface highlighting special effects, and the environmental feedback special effects into the multimodal synchronization engine according to the unified reference timeline for spatio-temporal feature alignment processing.
5. The AI character interaction method within the game platform according to claim 1, wherein, The step of determining the relationship link between the user behavior feature vector and at least one of the game scene, the task information, and the character information to generate a dynamic interaction strategy with interaction strategy priorities and parameter adjustment amplitudes includes: Construct a dynamic relationship graph including game scene objects, task target nodes, and character attribute nodes; Calculate the multi-dimensional association strength between the user behavior feature vector and at least one type of node in the dynamic relationship graph; Generate an initial set of interaction strategies based on the multi-dimensional association strength; Optimize the initial set of interaction strategies through a strategy conflict resolution model to generate a dynamic interaction strategy including strategy priority marks and parameter adjustment amplitude values.
6. The AI character interaction method in the game platform according to claim 5, wherein, The step of optimizing the initial set of interaction strategies through a strategy conflict resolution model to generate a dynamic interaction strategy including strategy priority marks and parameter adjustment amplitude values includes: Construct strategy conflict detection rules to detect logical contradictions and execution feasibility conflicts among the dialogue guidance strategy, the task assistance strategy, and the emotion feedback strategy in the initial set of interaction strategies, and generate corresponding conflict types; According to the detected conflict types, retrieve corresponding optimization strategies from a preset strategy optimization solution library to iteratively optimize the initial set of interaction strategies until all conflicts are eliminated; Sort and integrate the optimized initial set of interaction strategies according to the strategy priority marks and parameter adjustment amplitude values to generate the dynamic interaction strategy.
7. The AI character interaction method in the game platform according to claim 1, wherein, The step of displaying the multimodal interaction content to the user terminal includes: Deploy a pre-generated cache pool of the multimodal interaction content to an edge computing node connected to the user terminal and render potential interaction content in this edge computing node; Obtain the current network latency measurement value of the user terminal and determine the corresponding data transmission mode; Adjust the multimodal interaction content according to the data transmission mode and then send the adjusted multimodal interaction content to the user terminal for display.
8. The AI character interaction method in the game platform according to claim 1, characterized in that, The AI character interaction method further includes: Obtain the current game mode, where the game mode includes an online mode and a single-player mode; When the game mode is in the online mode, establish a network connection relationship among multiple user terminals, render the multi-modal interaction content displayed on the current user terminal from the perspective of the current user terminal, and synchronize the game status data to all user terminals participating in the online game in real time; When the game mode is in the single-player mode, disconnect the network connection with other user terminals, run the game logic on the local server of the user terminal, and store the corresponding game progress in the local storage medium.
9. An interaction device, characterized in that, The interaction device includes: a memory, a processor, and an AI character interaction program in the game platform stored on the memory and executable on the processor, and the AI character interaction program in the game platform is configured to implement the AI character interaction method in the game platform as described in any one of claims 1 to 8.
10. An interactive device, characterized in that, It includes the interaction device as described in claim 9.
Citation Information
Patent Citations
Game AI strategy decision model training method and device
CN111330279A
Game interaction control method and device, storage medium and electronic equipment
CN115581921A
Virtual character control method and device and electronic equipment
CN117883773A
Intelligent NPC system interacting with players
CN118807208A
Cited By
Whole-process engineering consultation project information interaction method and system
CN120578747A
Game development dynamic content generation method and system based on artificial intelligence
CN121327488A
Conversation processing method and device
CN122141256A
A dialog processing method and apparatus
CN122141256B