Method, device and equipment for interacting with AI characters in game platform

By generating dynamic interaction strategies and adjusting AI character behavior parameters, multimodal interactive content is rendered, solving the problem of the single form of AI character interaction in existing technologies, and improving the immersion of the game and user satisfaction.

CN120381658BActive Publication Date: 2025-12-16HANGZHOU KAIRONG NETWORK TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510750926.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-12-16
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing AI character interaction methods in games are monotonous in their presentation of interactive content, and their behavior is stiff and rigid. The interaction with users lacks realism and immersion, which exacerbates the sense of disconnect in the interactive experience, especially when multimodal interactive content is smoothly transmitted and presented.

Method used

By acquiring game scene, task information, and character information, a set of AI character behavior parameters is generated. User operation behavior and voice interaction data are collected, intent recognition processing is performed, user behavior feature vectors are generated, relationship links are determined and dynamic interaction strategies are generated, the set of behavior parameters is adjusted, and multimodal interactive content is rendered, including dialogue response mode, task guidance intensity, and emotional expression coefficient, and finally displayed on the user terminal.

Benefits of technology

It achieves realism and naturalness in AI character interaction, enhances user immersion and satisfaction in the game, solves the problem of smooth transmission of multimodal interactive content, and enhances the continuity of the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120381658B_ABST
    Figure CN120381658B_ABST
Patent Text Reader

Abstract

The application provides an AI role interaction method, an interaction device and an interaction equipment in a game platform, relates to the technical field of game interaction, and comprises the following steps: acquiring a current game scene, game task information and game character information, and generating a first behavior parameter set of an AI role; collecting current operation behavior data and voice interaction data of a user; performing intention recognition processing on the operation behavior data and the voice interaction data, generating a user behavior feature vector, determining a relationship link between the user behavior feature vector and at least one of game scene information, task information and character information, so as to generate a dynamic interaction strategy; generating a second behavior parameter set according to the dynamic interaction strategy, rendering and generating multi-modal interaction content matched with the second behavior parameter set; and displaying the multi-modal interaction content to a user terminal. The application can make the interaction of the AI role more real and natural, and can significantly improve the game immersion and satisfaction of the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of game interaction, in particular to an AI character interaction method, an interaction device and an interaction equipment in a game platform. BACKGROUND

[0002] The existing AI character interaction mode in the game has a single interaction content presentation form, the behavior performance is rigid and stereotyped, the interaction with the user lacks authenticity and immersion, and the game experience is seriously affected, especially in the smooth transmission and presentation of multi-modal interaction content, which exacerbates the fragmentation of the interaction experience. SUMMARY

[0003] The main purpose of the present application is to provide an AI character interaction method in a game platform, which aims to make the interaction of AI characters more real and natural, and can significantly improve the game immersion and satisfaction of users.

[0004] To achieve the above purpose, the present application provides an AI character interaction method in a game platform, which comprises:

[0005] Obtaining the current game scene, game task information and game character information, and generating a first behavior parameter set of the AI character;

[0006] Collecting the current operation behavior data and voice interaction data of the user;

[0007] Performing intent recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, the user behavior feature vector including user core intent, operation style feature and emotional tendency;

[0008] Determining the relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information to generate a dynamic interaction strategy with interaction strategy priority and parameter adjustment amplitude;

[0009] Adjusting the first behavior parameter set according to the dynamic interaction strategy to generate a second behavior parameter set, the first behavior parameter set including dialogue response mode, task guidance intensity and emotional expression coefficient;

[0010] Rendering and generating multi-modal interaction content matching the second behavior parameter set, the multi-modal interaction content including dialogue response mode, task guidance intensity and emotional expression coefficient;

[0011] Displaying the multi-modal interaction content to the user terminal.

[0012] Optionally, the intent recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector comprises:

[0013] inputting the voice interaction data after acoustic feature standardization processing into a speech recognition engine to convert the processed voice interaction data into a text sentence sequence with timestamp markers;

[0014] performing operation action semantic segmentation on the operation behavior data based on hot area distribution of a display interface of a user terminal, frequency intensity grading, and effectiveness markers of task progress states to generate a corresponding operation action sentence sequence;

[0015] constructing a space-time alignment model of the text sentence sequence and the operation action sentence sequence, performing timestamp dynamic matching, game object correlation mapping, and logic consistency detection to generate an intermediate feature set containing semantic keyword weight distribution, decision preference intensity indicators, and sentiment inclination indexes;

[0016] performing intent inference on the intermediate feature set based on a dynamically updated context state matrix to determine an intent probability distribution, to generate a user behavior feature vector corresponding to the intent probability distribution, the context state matrix containing a current task stage target, user historical behavior statistical features, and a scene object attribute relationship graph.

[0017] Optionally, the rendering and generating of the multi-modal interaction content matching the second behavior parameter set comprises:

[0018] based on a dialog response mode in the second behavior parameter set, calling a speech synthesis engine to generate speech waveform data with sentiment prosody features;

[0019] according to a task guidance intensity value in the second behavior parameter set, extracting a basic action primitive from a preset action library, performing action sequence mixing optimization on the basic action primitive through a physical simulation engine to generate three-dimensional role action data matching the task information, and generating an interface highlight special effect positively correlated with the guidance intensity value;

[0020] based on a sentiment expression coefficient in the second behavior parameter set, generating a facial micro-expression fusion weight, and controlling a particle special effect engine to generate an environmental feedback special effect matching the sentiment dimension based on the facial micro-expression fusion weight;

[0021] inputting the speech waveform data, the three-dimensional role action data, the interface highlight special effect, and the environmental feedback special effect into a multi-modal synchronization engine to generate multi-modal interaction content with space-time feature alignment.

[0022] Optionally, the inputting of the speech waveform data, the three-dimensional role action data, the interface highlight special effect, and the environmental feedback special effect into the multi-modal synchronization engine to generate multi-modal interaction content with space-time feature alignment comprises:

[0023] adding a first level timestamp mark to the voice waveform data according to a timing requirement of a dialogue response mode;

[0024] adding a second level timestamp mark to the three-dimensional character action data based on a task guidance intensity based physical simulation logic;

[0025] adding a third level timestamp mark to the environment feedback special effect according to a dynamic change curve of an emotional expression coefficient;

[0026] aligning the first level, the second level and the third level timestamp marks through a dynamic offset compensation algorithm to generate a unified reference time axis;

[0027] inputting the voice waveform data, the three-dimensional character action data, the interface highlight special effect and the environment feedback special effect into a multi-modal synchronization engine according to the unified reference time axis for spatiotemporal feature alignment processing.

[0028] Optionally, the determining of the relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information to generate a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment amplitude comprises:

[0029] constructing a dynamic relationship graph including game scene objects, task target nodes and character attribute nodes;

[0030] calculating a multi-dimensional association strength between the user behavior feature vector and at least one type of node in the dynamic relationship graph;

[0031] generating an initial interaction strategy set based on the multi-dimensional association strength;

[0032] optimizing the initial interaction strategy set through a strategy conflict resolution model to generate a dynamic interaction strategy including a strategy priority mark and a parameter adjustment amplitude value.

[0033] Optionally, the optimizing of the initial interaction strategy set through the strategy conflict resolution model to generate a dynamic interaction strategy including a strategy priority mark and a parameter adjustment amplitude value comprises:

[0034] constructing a strategy conflict detection rule to detect logical contradictions and execution feasibility conflicts among dialogue guidance strategies, task assistance strategies and emotional feedback strategies in the initial interaction strategy set, and generating corresponding conflict types;

[0035] according to the detected conflict types, calling corresponding optimization strategies from a preset strategy optimization scheme library to iteratively optimize the initial interaction strategy set until all conflicts are eliminated;

[0036] The initial interaction strategy set after optimization is sorted and integrated according to the strategy priority mark and the parameter adjustment amplitude value, and the dynamic interaction strategy is generated.

[0037] Optionally, the displaying the multi-modal interaction content to the user terminal comprises:

[0038] The pre-generated cache pool of the multi-modal interaction content is deployed to an edge computing node connected with the user terminal, and potential interaction content is rendered in the edge computing node;

[0039] A current network delay measurement value of the user terminal is obtained, and a corresponding data transmission mode is determined;

[0040] After adjusting the multi-modal interaction content according to the data transmission mode, the adjusted multi-modal interaction content is sent to the user terminal for display.

[0041] Optionally, the AI character interaction method further comprises:

[0042] A current game mode is obtained, and the game mode comprises an online mode and a single-player mode;

[0043] In a case where the game mode is in the online mode, a network connection relationship between a plurality of user terminals is established, and the multi-modal interaction content displayed by the current user terminal is rendered in a first perspective of the current user terminal, and game state data is synchronized in real time to all user terminals participating in online;

[0044] In a case where the game mode is in the single-player mode, the network connection with other user terminals is disconnected, and game logic is run in a local server of the user terminal, and corresponding game progress is stored in a local storage medium.

[0045] In addition, to achieve the above-mentioned purposes, the application further provides an interaction device, which comprises a memory, a processor, and an AI character interaction program stored in the memory and executable on the processor in a game platform, and the AI character interaction program in the game platform is configured to implement the AI character interaction method in the game platform as described above.

[0046] In addition, to achieve the above-mentioned purposes, the application further provides an interaction device, which comprises an interaction device as described above.

[0047] The embodiment of the application generates the first behavior parameter set of the AI role by acquiring the current game scene, game task information and game character information, collects the current operation behavior data and voice interaction data of the user, and performs intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, wherein the user behavior feature vector includes a user core intention, an operation style feature and an emotional tendency. After the user behavior feature vector is generated, the relationship link of the user behavior feature vector and at least one of the game scene, the task information and the character information is determined, a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment range is generated, the first behavior parameter set is adjusted according to the dynamic interaction strategy to generate a second behavior parameter set, wherein the first behavior parameter set includes a dialogue response mode, a task guidance intensity and an emotional expression coefficient, and the second behavior parameter set includes the dialogue response mode, the task guidance intensity and the emotional expression coefficient. The multi-modal interaction content matched with the second behavior parameter set is rendered and generated, wherein the multi-modal interaction content includes the dialogue response mode, the task guidance intensity and the emotional expression coefficient. Finally, the multi-modal interaction content is displayed to the user terminal to improve the user interaction smoothness and the interaction experience in the game platform. The embodiment can generate a dynamic interaction strategy and adjust the behavior parameters of the AI role according to the strategy, so as to render the multi-modal interaction content matched with the user behavior feature, make the interaction of the AI role more real and natural, and significantly improve the game immersion and satisfaction of the user. BRIEF DESCRIPTION OF DRAWINGS

[0048] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced in the following. Obviously, the accompanying drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0050] Figure 1 is a game platform AI role interaction method flowchart of an embodiment of the application;

[0051] Figure 2 is a game platform AI role interaction method flowchart of another embodiment of the application;

[0052] Figure 3 is a game platform AI role interaction method flowchart of another embodiment of the application;

[0053] Figure 4The flowchart of the AI role interaction method in the game platform of another embodiment of the present application is shown in Fig. 6;

[0054] Figure 5 The flowchart of the AI role interaction method in the game platform of another embodiment of the present application is shown in Fig. 6;

[0055] Figure 6 The flowchart of the AI role interaction method in the game platform of another embodiment of the present application is shown in Fig. 6;

[0056] Figure 7 The flowchart of the AI role interaction method in the game platform of another embodiment of the present application is shown in Fig. 6;

[0057] Figure 8 The flowchart of the AI role interaction method in the game platform of another embodiment of the present application is shown in Fig. 6.

[0058] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0059] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. The well-known modules, units and their connections, links, communications or operations are not shown or described in detail. And the described features, architectures or functions can be combined in any way in one or more embodiments. Those skilled in the art should understand that the following various embodiments are only used for illustration, but not for limiting the protection scope of the present application. It can also be easily understood that the modules or units or processing methods in the embodiments described herein and shown in the drawings can be combined and designed in various different configurations. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0060] In the following embodiments, the definition of various nouns or methods is generally based on the broad concept that can be implemented on the premise of the disclosed content in the embodiments, except in cases where it is logically impossible. Under such understanding, various specific lower specific definitions of the nouns or methods should be regarded as the invention content of the present application, and should not be regarded as the specific definition not disclosed in the specification, and should not be interpreted in a narrow sense or biased. Similarly, under the premise that the order of the steps in the method is flexible and variable, the specific lower specific definition of the broad concept of various nouns or methods is within the protection scope of the present application.

[0061] Due to the unsmooth user interaction process of the prior art game platform, the interactive experience is poor. The AI character interaction mode in the traditional game has obvious limitations: first, the understanding of the user's intention by the AI character is only based on simple keyword matching, which cannot accurately capture the core intention, operation style and emotional tendency implied in the user's operation behavior and voice interaction; second, the interaction strategy adopts a fixed template, lacking the ability to dynamically adjust according to the game scene, task progress and character attributes; third, the presentation form of the interaction content is single, and it is difficult to realize the space-time synchronization of multi-modal elements such as voice, action and special effects. These problems lead to the fact that the behavior of the AI character in the game process is stiff and stereotyped, and the interaction with the user lacks realism and immersion, which seriously affects the game experience. Especially when the network environment fluctuates, the prior art cannot guarantee the smooth transmission and presentation of multi-modal interaction content, greatly exacerbating the fragmentation of the interactive experience.

[0062] The main solution of the embodiment of the present application is: by acquiring the current game scene, game task information and game character information, generating a first behavior parameter set of the AI character, then collecting the user's current operation behavior data and voice interaction data, and performing intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, wherein the user behavior feature vector includes the user's core intention, operation style feature and emotional tendency. After the generation of the user behavior feature vector, determine the relationship link between the user behavior feature vector and at least one of the game scene, task information and character information, generate a dynamic interaction strategy with interaction strategy priority and parameter adjustment amplitude, then adjust the first behavior parameter set according to the dynamic interaction strategy to generate a second behavior parameter set, wherein the first behavior parameter set includes the dialogue response mode, task guidance intensity and emotional expression coefficient, and the second behavior parameter set includes the dialogue response mode, task guidance intensity and emotional expression coefficient. Then render and generate multi-modal interaction content matching the second behavior parameter set, wherein the multi-modal interaction content includes the dialogue response mode, task guidance intensity and emotional expression coefficient. Finally, display the multi-modal interaction content to the user terminal.

[0063] In this embodiment, for ease of description, the following describes the interactive device as the execution subject.

[0064] The present application provides a solution to improve the smoothness of user interaction and interactive experience in the game platform. The embodiment can generate a dynamic interaction strategy and adjust the behavior parameters of the AI character according to the strategy, thereby rendering multi-modal interaction content matching the user behavior features, making the interaction of the AI character more realistic and natural, and significantly improving the user's game immersion and satisfaction.

[0065] To this end, the application provides an AI role interaction method in a game platform. It can be understood that the interaction device is provided with an interaction device for storing and executing the following method, and the interaction device can be realized by a main controller such as MCU (Micro controller Unit), DSP (Digital Signal Process), FPGA (Field Programmable Gate Array), SOC (System On Chip) and the like.

[0066] In the prior art, the game platform user interaction process generally has the problems of high operation delay and low intention recognition accuracy. The traditional system relies on a preset script to drive the AI role behavior, and cannot adjust the interaction strategy according to the real-time game scene and user dynamics. In a typical scenario, when the user completes a multi-player copy task, the timing of the voice instruction and the interface operation is misaligned, resulting in that the task guide provided by the AI role does not match the current battle stage.

[0067] In order to solve the above problems, it is necessary to break through the static strategy framework and build a dynamic and adaptive multi-modal interaction mechanism. It is found in the design process that a single data source analysis is difficult to capture the user's composite intention, and it is necessary to fuse voice, operation behavior and scene context for joint modeling.

[0068] Based on the above content, with reference to Figure 1 In an embodiment of the application, the AI role interaction method in the game platform comprises steps S100-S700, wherein:

[0069] S100, obtaining the current game scene, game task information and game character information, and generating a first behavior parameter set of the AI role;

[0070] S200, collecting the current operation behavior data and voice interaction data of the user;

[0071] S300, performing intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector;

[0072] S400, determining the relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information to generate a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment amplitude;

[0073] S500, adjusting the first behavior parameter set according to the dynamic interaction strategy to generate a second behavior parameter set, wherein the first behavior parameter set includes a dialogue response mode, a task guide intensity and an emotional expression coefficient;

[0074] S600, render and generate multi-modal interactive content matching the second set of behavior parameters, the multi-modal interactive content including a dialogue response mode, a task guidance intensity, and an emotional expression coefficient;

[0075] S700, display the multi-modal interactive content to the user terminal.

[0076] The user behavior feature vector includes a user core intent, an operation style feature, and an emotional tendency. The user behavior feature vector refers to a user interaction feature set extracted through multi-dimensional data analysis, which can be realized by acoustic feature standardization processing and operation action semantic segmentation technology, and is used to accurately represent the user core intent and emotional state. The dynamic interaction strategy refers to an interaction scheme generated based on real-time context, which can be realized by constructing a dynamic relationship graph to calculate the multi-dimensional correlation strength, and is used to guide the dynamic adjustment of the AI character behavior parameters. The multi-modal interactive content refers to a composite output that integrates voice, action, and visual special effects, which can be realized by using a multi-modal synchronization engine to align the time stamp, ensuring that the interactive content of different sensory channels maintains spatiotemporal consistency.

[0077] The interaction device obtains scene terrain data, task target list, and NPC attribute parameters through a game engine interface to initialize the basic behavior mode of the AI character. For example, when the user performs a rescue task in a shooting game, the voice command "shield left" and the operation data of quickly switching weapons are synchronously collected. The user frequently clicks on the left side of the shelter area through hot area distribution analysis, and the tactical instruction keywords are extracted in combination with voice recognition to generate a feature vector containing tactical deployment intent and aggressive operation style. The interaction device correlates the feature vector with the task progress stage and the enemy firepower distribution to dynamically improve the left side defense frequency and firepower suppression intensity of the AI teammate. The adjusted behavior parameters drive the generation of voice prompts "blocking the left channel", the animation of the character quickly moving to the shelter position, and the highlight boundary special effect of the left area, forming a coordinated multi-modal feedback.

[0078] Compared with the prior art, the traditional method uses a fixed priority strategy to process user input, which is prone to strategy conflicts when the voice command and operation behavior point to different targets. The embodiment realizes the fusion and analysis of multi-source data through a spatiotemporal alignment model, and quantifies the correlation degree of user intent and game elements in combination with a dynamic relationship graph, effectively solving the problem of collaborative processing of cross-modal instructions. In an open-world exploration scenario, the user asks "treasure location" while operating the character to move to the map marker, and the interaction device can accurately identify the redundant guidance demand and automatically reduce the task prompt frequency.

[0079] By means of the above technical means, the embodiment realizes the context perception and adaptive interaction of AI characters in a game environment. In a real-time strategy game, when a user simultaneously performs voice command and multi-unit micro-operation, the interaction device can accurately identify the core tactical intention and dynamically adjust the autonomous behavior parameters of AI units. In a role-playing task, the facial expression and voice tone of NPCs are automatically matched according to the emotional tendency of the user, thereby enhancing the immersion and naturalness of the interaction process.

[0080] The embodiment generates a first behavior parameter set of an AI character by acquiring current game scene, game task information and game character information, collects user's current operation behavior data and voice interaction data, and performs intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, wherein the user behavior feature vector includes user core intention, operation style feature and emotional tendency. After the generation of the user behavior feature vector, the relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information is determined, a dynamic interaction strategy with interaction strategy priority and parameter adjustment amplitude is generated, and the first behavior parameter set is adjusted according to the dynamic interaction strategy to generate a second behavior parameter set. The first behavior parameter set includes dialogue response mode, task guidance intensity and emotional expression coefficient, and the second behavior parameter set includes dialogue response mode, task guidance intensity and emotional expression coefficient. The second behavior parameter set is rendered and matched with the second behavior parameter set to generate multi-modal interaction content, and the multi-modal interaction content includes dialogue response mode, task guidance intensity and emotional expression coefficient. Finally, the multi-modal interaction content is displayed to the user terminal to improve the user interaction smoothness and interaction experience in the game platform. The embodiment can generate a dynamic interaction strategy and adjust the behavior parameters of the AI character according to the strategy, thereby rendering multi-modal interaction content matched with the user behavior feature, making the interaction of the AI character more realistic and natural, and significantly improving the game immersion and satisfaction of the user.

[0081] Optionally, referring to Figure 2 Another embodiment of the present application provides an AI character interaction method in a game platform, based on the above Figure 1 As shown in the embodiment, the intention recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector includes steps S310-S340, wherein:

[0082] S310, after the acoustic feature standardization processing of the voice interaction data, the processed voice interaction data is input into a voice recognition engine to convert the processed voice interaction data into a text sentence sequence with a timestamp mark;

[0083] S320, performing operation action semantic segmentation on the operation behavior data based on the hot area distribution of the display interface of the user terminal, the frequency force degree classification, and the effectiveness mark of the task progress state, to generate a corresponding operation action sentence sequence;

[0084] S330, constructing a space-time alignment model of the text sentence sequence and the operation action sentence sequence, performing timestamp dynamic matching, game object correlation mapping, and logic consistency detection, to generate an intermediate feature set containing semantic keyword weight distribution, decision preference intensity index, and sentiment tendency index;

[0085] S340, performing intent inference on the intermediate feature set based on a dynamically updated context state matrix, determining an intent probability distribution, to generate a user behavior feature vector corresponding to the intent probability distribution.

[0086] The acoustic feature standardization processing refers to noise reduction, tone unification, and dialect adaptation processing of the voice data, and can also use a deep neural network-based noise reduction algorithm, a mel-frequency cepstral coefficient analysis, and a dialect voice library matching algorithm to achieve, for eliminating environmental interference and improving the recognition compatibility of different accents. The operation action semantic segmentation refers to structured analysis of operation behavior according to the hot area distribution of user operation, operation frequency force, and task progress effectiveness, which can be realized by an interface hot area tracking module, a pressure sensor data acquisition, and a task state mark interaction device, for converting discrete operation behavior into sequence data with semantic information. The space-time alignment model refers to establishing the correlation between voice text and operation action in time and space dimensions, which can be realized by a dynamic time warping algorithm, a game object recognition engine, and a logic rule checker, for ensuring the consistency of voice instructions and operation behavior in time sequence and logic. The context state matrix refers to a dynamic data set integrating the current task target, user historical behavior mode, and scene object attribute, which can be realized by a task state tracker, a user behavior log analyzer, and a scene attribute database, for providing real-time context support for intent inference.

[0087] The voice data is converted into a timestamped text sequence after noise reduction, tone normalization and dialect adaptation, eliminating the influence of environmental noise and accent difference on recognition accuracy. The user attention area is determined according to the interface hot area distribution, the operation intensity is quantified by combining the operation frequency and intensity, and the invalid operation is filtered by the task progress effectiveness marker, and the operation behavior data is segmented into action sentences with semantics. The text sequence and the action sentence are dynamically matched and aligned in time by the timestamp, the mapping relationship between the voice instruction and the operation target is established by using the game object recognition engine, the contradictory data is excluded by the logic rule checker, and the intermediate features including the keyword weight, the decision preference and the emotional tendency are formed. The current task stage target, the user historical behavior statistical features and the scene object attribute relationship are integrated by the dynamically updated context matrix, the intention distribution is calculated by the probability model, and finally the behavior feature vector reflecting the real-time intention of the user is generated.

[0088] Compared with the prior art, the traditional method usually processes voice and operation data separately, resulting in isolated intention recognition and lack of context association. The voice recognition in the prior art does not consider dialect difference and environmental noise interference, the operation behavior analysis does not combine task progress effectiveness and interface hot area distribution, and the space-time alignment only uses simple timestamp matching without establishing game object association mapping. The embodiment realizes multi-source data preprocessing by acoustic standardization and operation semantic segmentation, integrates time sequence and logical relationship by constructing a space-time alignment model, and realizes intention inference by combining a dynamic context matrix, solving the intention deviation problem caused by isolated data processing.

[0089] Through the above technical means, the embodiment can effectively integrate multi-dimensional data of user voice and operation behavior, eliminate the influence of environmental interference and operation noise on intention recognition, and accurately extract semantic keyword weight and decision preference features. Through intention probability inference supported by dynamic context, the generated user behavior feature vector accurately reflects the core intention and behavior pattern of the user in the current game scene, providing a reliable data basis for subsequent interaction strategy generation, and significantly improving the response accuracy and scene adaptability of the game AI character to user intention.

[0090] Optionally, referring to Figure 3 , another embodiment of the application provides an AI character interaction method in a game platform, based on the above Figure 1 The rendering and generation of the multi-modal interaction content matched with the second behavior parameter set includes steps S610-S640, wherein:

[0091] S610, based on the dialogue response mode in the second behavior parameter set, a voice synthesis engine is called to generate voice waveform data with emotional rhythm characteristics;

[0092] S620, extracting a basic action primitive from a preset action library according to the task guidance intensity value in the second behavior parameter set, performing action sequence mixing optimization on the basic action primitive through a physical simulation engine, generating three-dimensional role action data matching the task information, and generating an interface highlight special effect positively correlated with the guidance intensity value;

[0093] S630, generating a facial micro-expression fusion weight based on the emotional expression coefficient in the second behavior parameter set, and controlling a particle special effect engine to generate an environmental feedback special effect matching the emotional dimension based on the facial micro-expression fusion weight;

[0094] S640, inputting the voice waveform data, the three-dimensional role action data, the interface highlight special effect, and the environmental feedback special effect into a multi-modal synchronous engine to generate multi-modal interaction content with spatiotemporal feature alignment.

[0095] The dialogue response mode refers to voice output logic matched according to user intent, which can be implemented by using a voice synthesis model based on emotional classification, and its function is to ensure semantic consistency between voice content and current user intent. The task guidance intensity value refers to a quantized task promotion parameter, which can be calculated by using a sliding window to count the relationship between task completion degree and operation frequency, and its function is to dynamically adjust the amplitude of role action and the intensity of interface prompt. The emotional expression coefficient refers to a weight distribution parameter of multi-dimensional emotional feedback, which can be generated by using an emotional dimension mapping matrix combined with real-time emotional recognition results, and its function is to coordinate the synchronous expression of facial expression and environmental special effect. The multi-modal synchronous engine refers to a spatiotemporal alignment data processing module, which can be implemented by using a dynamic time warping algorithm combined with an event-driven mechanism, and its function is to eliminate the time offset and logical conflict of different modal data.

[0096] In the process of generating multi-modal interaction content, the voice synthesis engine adjusts the fundamental frequency and speech rate parameters according to the emotional classification results in the dialogue response mode to generate waveform data matching the user intent. The physical simulation engine optimizes the trajectory of the extracted basic action primitive through inverse kinematics algorithm to generate a three-dimensional action sequence conforming to the task guidance intensity, and linearly adjusts the light-emitting frequency of the interface elements according to the guidance intensity value. The facial micro-expression fusion weight superimposes multiple basic expressions according to the emotional coefficient ratio through hybrid deformation technology, and the particle special effect engine adjusts the particle motion trajectory and color gradient mode according to the emotional dimension parameters. The multi-modal synchronous engine captures the output signals of each module through the event-driven mechanism, aligns the time axes of voice waveform, role action, and environmental special effect by using the dynamic time warping algorithm, and ensures that all modal data reach the synchronous output state at the interaction triggering moment.

[0097] Compared with the prior art, the traditional method adopts independent generation of each modal data and then simple superposition, which is easy to cause the disconnection of voice and action, and the incoordination of special effect and expression. Through the establishment of the parameter linkage mechanism and the multi-modal synchronization engine, the embodiment establishes the space-time correlation between elements in the data generation stage, realizes the dynamic binding of action primitives and task guidance intensity, the dimensional mapping of emotional coefficient and environmental special effect, and eliminates the response delay of cross-modal through a unified time reference.

[0098] Through the above technical means, the embodiment solves the space-time asynchronization problem in the multi-modal interactive content generation process, ensures the logical consistency of voice guidance and character action in the task promotion process, realizes the dynamic matching of facial expression change and environmental special effect, and effectively improves the overall perception fluency of the user on the interactive content. The application of the multi-modal synchronization engine makes different interactive elements form a synergistic effect in the triggering time and performance intensity, avoids the time misalignment phenomenon of visual cues and voice feedback in the traditional method, and enhances the effectiveness of task guidance and the immersion of emotional expression.

[0099] Optionally, referring to Figure 4 , a further embodiment of the application provides an AI character interaction method in a game platform, based on the above Figure 3 , the voice waveform data, the three-dimensional character action data, the interface highlight special effect and the environmental feedback special effect are input into the multi-modal synchronization engine to generate multi-modal interactive content with space-time feature alignment, including steps S641-S645, wherein:

[0100] S641, according to the time sequence requirement of the dialogue response mode, adding a first level timestamp mark to the voice waveform data;

[0101] S642, based on the physical simulation logic of the task guidance intensity, adding a second level timestamp mark to the three-dimensional character action data;

[0102] S643, according to the dynamic change curve of the emotional expression coefficient, adding a third level timestamp mark to the environmental feedback special effect;

[0103] S644, aligning the first, second and third level timestamp marks through a dynamic offset compensation algorithm to generate a unified reference time axis;

[0104] S645, according to the unified reference time axis, inputting the voice waveform data, the three-dimensional character action data, the interface highlight special effect and the environmental feedback special effect into the multi-modal synchronization engine for space-time feature alignment processing.

[0105] The first level timestamp mark refers to a synchronous mark set based on the timing characteristics of the voice output of the dialogue response mode, which can be realized by using the voice segmentation processing technology combined with the voice activity detection algorithm, and is used to ensure that the natural rhythm of the voice waveform data matches the user operation in real time. The second level timestamp mark refers to the time coding of the role action sequence according to the physical simulation logic, which can be generated by using the rigid body dynamics model combined with the motion trajectory interpolation algorithm, so that the three-dimensional role action conforms to the physical motion law and is synchronized with the change of the task guide intensity. The third level timestamp mark refers to the special effect trigger mark generated according to the emotional expression curve, which can be realized by using the time sequence prediction model combined with the emotional dimension mapping rule, to ensure that the particle emission frequency of the environmental feedback special effect is consistent with the change trend of the emotional tendency. The dynamic offset compensation algorithm refers to a synchronous correction method for eliminating the time error of the multi-modal data stream, which can be realized by using the clock drift estimation model combined with the adaptive delay compensation mechanism, by monitoring the time offset of each data stream in real time and dynamically adjusting the time reference, a unified synchronous time axis is generated.

[0106] When the voice waveform data is added with the first level timestamp mark, the voice segment boundary is divided by using the voice endpoint detection technology, and the time mark interval is set according to the rhythm requirement of the dialogue response mode, for example, an extension mark is set at the end of a question sentence. When the three-dimensional role action data is added with the second level timestamp mark, the key action frame is time-coded according to the skeletal motion trajectory calculated by the physical engine, for example, different time marks are set for the three phases of take-off, hovering and landing when the role performs a jumping action. When the environmental feedback special effect is added with the third level timestamp mark, the special effect intensity adjustment mark is set at the inflection point of the curve according to the numerical change rate of the emotional expression coefficient, for example, a particle burst mark is set when the emotional coefficient suddenly increases. The dynamic offset compensation algorithm calculates the time deviation of each stream by monitoring the transmission delay of the voice, action and special effect three data streams in real time, for example, when the voice stream is 5 milliseconds faster than the action stream, the time mark of the action stream is compensated forward, and a unified time reference axis is generated. The multi-modal synchronous engine performs frame synchronization processing on each data stream according to the time axis, for example, the voice waveform playback, role action execution and interface special effect rendering are triggered at the same time at a certain time point.

[0107] Compared with the prior art, the traditional multi-modal synchronization method usually adopts a single timestamp marking system, which cannot adapt to the generation mechanism differences of different data streams, for example, speech data adopts fixed frame rate marking while action data relies on physical simulation calculation, which is easy to cause the lips to be out of sync with the speech. In the prior art, environmental special effects usually adopt preset trigger conditions, which cannot dynamically adapt to changes in emotional expression, for example, special effects may be delayed when the player's emotions are intense. The embodiment establishes a hierarchical timestamp system, sets differentiated marking rules for the generation principles of different modal data, and combines a dynamic compensation mechanism to eliminate cross-modal time errors, for example, when physical simulation calculation delay causes action data to lag, the presentation timing of other data streams is automatically adjusted.

[0108] Through the above technical means, the embodiment solves the problem of space-time asynchronization of multi-modal data caused by differences in generation mechanisms, accurately matches the speech output rhythm with the character lip shape changes, keeps the task guide actions and interface prompt special effects in coordination, and adjusts the environmental particle special effect intensity in real time according to the emotional changes. The technical defects that may occur in the traditional method, such as the character starting to act after the speech is finished and the emotional special effects disappearing too early to affect the immersion, are avoided, and accurate synchronization presentation of multi-dimensional interactive elements is realized.

[0109] Optionally, referring to Figure 5 , the application also provides an AI character interaction method in a game platform, based on the above Figure 1 The embodiment shown, the determination of the relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information is to generate a dynamic interaction strategy with interaction strategy priority and parameter adjustment amplitude, including steps S410-S440, wherein:

[0110] S410, a dynamic relationship graph containing game scene objects, task target nodes and character attribute nodes is constructed;

[0111] S420, the multi-dimensional association strength of the user behavior feature vector and at least one type of node in the dynamic relationship graph is calculated;

[0112] S430, an initial interaction strategy set is generated based on the multi-dimensional association strength;

[0113] S440, the initial interaction strategy set is optimized by a strategy conflict resolution model to generate a dynamic interaction strategy containing strategy priority labels and parameter adjustment amplitude values.

[0114] The dynamic relationship graph refers to a data structure in which interactive objects, task targets and character attributes in a game scene are associated in the form of nodes, which can be implemented by using a graph database or a knowledge graph technology, and is used for integrating multi-dimensional information in a game environment in real time. The multi-dimensional correlation strength refers to a similarity index between a user behavior feature and a graph node calculated by using a vector space model, which can be implemented by using a cosine similarity algorithm or a neural network embedding technology, and is used for quantifying the correlation degree of a user intention and a game element. The initial interaction strategy set refers to a candidate strategy combination generated based on the correlation strength, which can be implemented by using a rule engine or a probability model, and is used for covering interaction requirements in different scenes. The strategy conflict resolution model refers to an optimization module for detecting logical contradictions between strategies, which can be implemented by using a constraint satisfaction algorithm or a reinforcement learning model, and is used for eliminating conflicts in the strategy execution process.

[0115] The dynamic relationship graph is constructed to integrate game scene, task target and character attribute information, thereby providing a structured data basis for subsequent correlation strength calculation. When calculating the multi-dimensional correlation strength between a user behavior feature vector and a graph node, a vector similarity algorithm is used to quantify the correlation degree of a user intention and a game element, thereby generating an initial strategy set covering different interaction dimensions. Logical contradictions between dialogue guidance, task assistance and emotional feedback strategies are identified by using a conflict detection rule, and a conflict resolution scheme in an optimized strategy library is called to perform iterative processing. Finally, the optimized strategies are sorted and integrated according to priority and adjustment amplitude, thereby forming a conflict-free dynamic interaction strategy.

[0116] Compared with the prior art, the traditional method usually generates a strategy based on single-dimensional user behavior data, lacks dynamic correlation analysis of game scenes, tasks and character attributes, and is prone to execution conflicts when multiple strategies are executed in parallel. The embodiment realizes multi-dimensional information integration by constructing a dynamic relationship graph, and optimizes the strategies by using a conflict resolution model, thereby effectively coordinating the execution logic between different strategies.

[0117] By using the above technical means, the embodiment solves the problem of logical conflicts when multiple strategies are executed in parallel, avoids interaction interruption or behavior confusion caused by strategy contradictions, and improves the fluency of the interaction process and the consistency of strategy execution.

[0118] Optionally, referring to Figure 6 , another embodiment of the present application provides an AI character interaction method in a game platform, based on the above Figure 5 embodiment, the initial interaction strategy set is optimized by using a strategy conflict resolution model to generate a dynamic interaction strategy containing a strategy priority label and a parameter adjustment amplitude value, including steps S441-S443, wherein:

[0119] S441, construct a strategy conflict detection rule to detect logical contradictions and execution feasibility conflicts among the dialogue guiding strategy, the task assisting strategy, and the emotional feedback strategy in the initial interaction strategy set, and generate a corresponding conflict type;

[0120] S442, according to the detected conflict type, call the corresponding optimization strategy from the preset strategy optimization scheme library to perform iterative optimization processing on the initial interaction strategy set until all conflicts are eliminated.

[0121] S443, sort and integrate the initial interaction strategy set after optimization according to the strategy priority label and the parameter adjustment amplitude value to generate the dynamic interaction strategy.

[0122] The strategy conflict detection rule refers to a verification mechanism for identifying logical contradictions and execution feasibility conflicts between different strategies, which can be implemented by semantic logic verification, resource occupation simulation, or execution timing analysis, and is used to detect implicit conflicts between strategies. The preset strategy optimization scheme library refers to a database that stores the mapping relationship between conflict types and optimization strategies, which can be implemented by a hash index structure based on conflict tags or a graph neural network matching algorithm, and is used to quickly locate the applicable optimization strategy. The strategy priority label refers to metadata for identifying the execution order of the strategy, which can be implemented by a weight scoring mechanism or a dependency graph, and is used to determine the priority order when multiple strategies are executed in parallel. The parameter adjustment amplitude value refers to a quantitative indicator of the modification degree of the strategy parameters, which can be implemented by a gradient descent algorithm or a fuzzy logic controller, and is used to control the boundary range of strategy adjustment.

[0123] By constructing the strategy conflict detection rule, logical contradictions are detected among the dialogue guiding strategy, the task assisting strategy, and the emotional feedback strategy in the initial interaction strategy set. For example, when the task assisting strategy requires forced progress of the task progress, and the emotional feedback strategy detects negative emotions of the user, an execution feasibility conflict is triggered. After detecting the conflict type, the optimization strategy corresponding to the conflict type is called from the preset strategy optimization scheme library. For example, for the conflict between task promotion and emotional soothing, an optimization strategy combining task target splitting and emotional compensation mechanism is adopted. Through iterative optimization processing, such as performing conflict detection again after the first optimization, it is ensured that all conflicts are eliminated. The optimized strategy set is assigned a priority label, such as setting the task assisting strategy as high priority and the emotional feedback strategy as dynamic adjustment priority, while limiting the modification range of the strategy parameters according to the parameter adjustment amplitude value, and finally generating a conflict-free dynamic interaction strategy.

[0124] Compared with the prior art, the traditional method adopts a single strategy optimization or manually sets priorities, and cannot interactively solve the implicit conflicts in the collaborative execution of multiple strategies, for example, only covering contradictions by weight adjustment, resulting in logical discontinuity in the interaction process. The prior art lacks a precise matching mechanism for the optimization strategy of conflict types, and often uses fixed priority sorting, resulting in strategy rigidity. The embodiment realizes the root cause elimination of strategy conflicts through a hierarchical conflict detection and iterative optimization mechanism, while preserving the integrity and execution flexibility of the strategy system.

[0125] Through the above technical means, the embodiment solves the interaction process lag problem caused by logical contradictions in the collaborative execution of multiple strategies, avoids parameter loss of control in the strategy adjustment process through conflict type identification and targeted optimization strategy calling, ensures that the dynamically generated interaction strategy maintains the balance between task promotion efficiency and emotional feedback while eliminating internal conflicts, and improves the interaction continuity of users in complex game scenarios.

[0126] Optionally, referring to Figure 7 , another embodiment of the application provides an AI character interaction method in a game platform, based on the above Figure 1 The embodiment shown in the embodiment, the display of the multi-modal interaction content to the user terminal includes steps S710-S730, wherein:

[0127] S710, deploy the pre-generated cache pool of the multi-modal interaction content to the edge computing node connected with the user terminal, and render the potential interaction content in the edge computing node;

[0128] S720, obtain the current network delay measurement value of the user terminal, and determine the corresponding data transmission mode;

[0129] S730, adjust the multi-modal interaction content according to the data transmission mode, and then send the adjusted multi-modal interaction content to the user terminal for display.

[0130] A dynamic code rate adjustment mechanism is established, and the following transmission modes are adaptively selected according to the network delay measurement value: when the delay is less than 50 milliseconds, the complete multi-modal data stream is transmitted; when the delay is between 50 milliseconds and 150 milliseconds, the progressive loading strategy is enabled; when the delay is greater than or equal to 150 milliseconds, switch to the simplified interaction protocol, and only keep the key semantic data.

[0131] The pre-generated cache pool refers to the pre-stored rendered interactive content data blocks in the edge computing node, which can be implemented by using a distributed storage technology combined with a rendering task queue management, and is used to reduce the computing pressure and data transmission volume of real-time rendering. The edge computing node refers to a server cluster deployed in the vicinity of the user terminal, which can be implemented through a content distribution network architecture, and is used to shorten the data transmission path and reduce the network delay. The network delay measurement value refers to the communication response time index between the user terminal and the server, which can be implemented by using a round-trip time measurement algorithm combined with a data packet loss rate detection, and is used to evaluate the current network transmission quality. The data transmission mode refers to a combination of data transmission strategies dynamically adjusted according to the network state, which can be implemented by using a protocol stack parameter configuration combined with a data compression algorithm, and is used to adapt to the transmission requirements under different network conditions. The dynamic code rate adjustment mechanism refers to a decision model for automatically switching transmission strategies based on the network delay volume, which can be implemented by using a multi-level threshold trigger rule combined with a priority queue scheduling, and is used to balance the data integrity and transmission efficiency.

[0132] The pre-generated cache pool generates potential interactive content in advance through the local rendering capability of the edge computing node, avoiding the high delay problem caused by remote rendering in the cloud. The network delay measurement value is monitored in real time and classified into different levels, triggering the corresponding transmission mode: in the low delay scenario, the complete data stream is transmitted to ensure the details of the interactive content; in the medium delay scenario, the core interactive elements are preferentially transmitted and the auxiliary content is gradually loaded; in the high delay scenario, only the key semantic data is retained to maintain the basic interactive logic. The dynamic code rate adjustment mechanism divides the network state interval through multi-level thresholds, ensures that the transmission strategy matches the network fluctuation in real time, and reduces the dependence on the cloud data by combining the local cache of the edge node.

[0133] Compared with the prior art, the traditional method uses a fixed code rate transmission mode, which cannot cope with the delay changes caused by network fluctuations, and is prone to cause data congestion or content loss. In the prior art, the rendering of interactive content usually relies on the cloud server, resulting in a long transmission path and difficulty in real-time response. The embodiment realizes localized rendering and caching by deploying edge computing nodes, and combines a dynamic code rate adjustment mechanism, effectively shortening the data transmission path and adapting to network state changes.

[0134] Through the above technical means, the embodiment solves the problem of interactive content transmission delay or lag caused by network delay, and ensures the continuity and real-time performance of multi-modal interactive content under different network conditions. Through edge node pre-rendering and dynamic transmission strategy, the data transmission load between the cloud and the terminal is reduced, and the risk of network congestion in the fixed code rate mode is avoided. When the network state fluctuates, the data volume of the transmitted content is adaptively adjusted, which not only maintains the availability of the core interactive function, but also optimizes the user experience in the high bandwidth scenario.

[0135] Optionally, with reference toFigure 8 Another embodiment of the present application provides an AI character interaction method in a game platform, based on the above Figure 1 As shown in the embodiment, the AI character interaction method further comprises steps S800-S1000, wherein:

[0136] S800, acquire the current game mode, the game mode including online mode and single-player mode;

[0137] S900, in the case where the game mode is in the online mode, establish a network connection relationship between multiple user terminals, and render the multi-modal interaction content displayed by the current user terminal from a first perspective of the current user terminal, and synchronize the game state data to all user terminals participating in online in real time;

[0138] S1000, in the case where the game mode is in the single-player mode, disconnect the network connection with other user terminals, and run the game logic in the local server of the user terminal, and store the corresponding game progress to the local storage medium.

[0139] Among them, the game mode refers to the state classification of the running environment, which can be realized by user active selection or interactive device automatic detection, used to trigger different network connection strategies and data processing mechanisms. The network connection relationship refers to the communication link topology structure between terminals, which can be established by P2P protocol or central server architecture, used to ensure the data interaction needs between multiple users in online mode. The first perspective rendering refers to the display mode centered on the current user operation interface, which can be realized by perspective locking algorithm, used to ensure the primary and secondary hierarchical relationship of the interaction content. The local server refers to the logic operation unit deployed in the user terminal, which can be realized by lightweight virtualization technology, used to independently process the game running and storage needs in single-player mode.

[0140] Among them, when the interactive device detects that the user enters the online mode, the network connection module automatically triggers to build the communication channel between multiple terminals. For example, after confirming the online state of each terminal by the heartbeat packet detection mechanism, the multi-modal interaction content of the current user is rendered by priority using the distributed synchronization protocol, and the game state data is compressed and encapsulated and then transmitted to other participating terminals. In this process, the perspective locking algorithm continuously tracks the user operation interface to ensure that the three-dimensional scene rendering and interaction feedback are presented from the first perspective. When switching to single-player mode, the network connection module actively releases the communication resources, and the local server takes over the game logic operation, for example, maintains the running state of the game progress by memory resident technology, and at the same time, encrypts the key progress data and writes it to the local storage medium.

[0141] Compared with the prior art, the traditional scheme needs to manually adjust the network configuration when switching modes and cannot maintain data consistency, while the embodiment realizes real-time synchronization of multi-terminal data in online mode and offline autonomous operation in single-machine mode through dynamic network topology reconstruction and localization processing mechanism. The online mode in the prior art relies on fixed server architecture, resulting in high delay, and the embodiment uses a distributed communication protocol to reduce data transmission delay; the single-machine mode relies on external storage, which has security risks, and the embodiment enhances data security through local encrypted storage.

[0142] Through the above technical means, the embodiment solves the problem of interactive confusion caused by different synchronization of multi-user perspectives in online mode, avoids resource waste caused by network connection redundancy in single-machine mode, and at the same time, guarantees the integrity and recoverability of game progress through the localization storage mechanism. The real-time synchronization mechanism in online mode ensures the consistency of the state of all terminals, and the independent running logic in single-machine mode improves the operation response speed, and the seamless switching of the two modes optimizes the overall interactive experience.

[0143] The application further provides an interactive device, which comprises a memory, a processor, and an AI character interaction program in a game platform stored in the memory and executable on the processor.

[0144] It should be noted that, since the interactive device is based on the above-mentioned AI character interaction method in the game platform, the embodiments of the interactive device include all the technical solutions of all the embodiments of the AI character interaction method in the game platform, and the technical effects achieved are also exactly the same, which will not be repeated here.

[0145] The application further provides an interactive device, which comprises a memory, a processor, and an AI character interaction program in a game platform stored in the memory and executable on the processor.

[0146] It should be noted that, since the interactive device is based on the above-mentioned AI character interaction method in the game platform, the embodiments of the interactive device include all the technical solutions of all the embodiments of the AI character interaction method in the game platform, and the technical effects achieved are also exactly the same, which will not be repeated here.

[0147] It should be noted that, in this document, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element preceded by "comprises... a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.

[0148] The above-mentioned embodiment numbers of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0149] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and the necessary general hardware platform, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) as described above, and includes a number of instructions for making a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) execute the methods described in the various embodiments of the present application.

[0150] The above is only the preferred embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent flow transformation made by using the content of the specification and drawings, or directly or indirectly applied to other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An AI character interaction method in a game platform, characterized in that, The AI character interaction method comprises: obtaining a current game scene, game task information and game character information, and generating a first behavior parameter set of an AI character; collecting current operation behavior data and voice interaction data of a user; performing intent recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector, the user behavior feature vector comprising a user core intent, an operation style feature and an emotional tendency; determining a relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information to generate a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment amplitude; adjusting the first behavior parameter set according to the dynamic interaction strategy to generate a second behavior parameter set, the first behavior parameter set comprising a dialogue response mode, a task guidance intensity and an emotional expression coefficient; rendering and generating multi-modal interaction content matching the second behavior parameter set, the multi-modal interaction content comprising a dialogue response mode, a task guidance intensity and an emotional expression coefficient; displaying the multi-modal interaction content to a user terminal; The determination of the relationship link between the user behavior feature vector and at least one of the game scene, the task information and the character information to generate a dynamic interaction strategy with an interaction strategy priority and a parameter adjustment amplitude comprises: constructing a dynamic relationship graph comprising game scene objects, task target nodes and character attribute nodes; calculating a multi-dimensional association strength between the user behavior feature vector and at least one type of node in the dynamic relationship graph; generating an initial interaction strategy set based on the multi-dimensional association strength; optimizing the initial interaction strategy set through a strategy conflict resolution model to generate a dynamic interaction strategy comprising a strategy priority label and a parameter adjustment amplitude value.

2. The method of Claim 1, wherein, The intent recognition processing on the operation behavior data and voice interaction data to generate a user behavior feature vector comprises: inputting the voice interaction data after acoustic feature standardization processing into a speech recognition engine to convert the processed voice interaction data into a text sentence sequence with a timestamp label; performing operation action semantic segmentation on the operation behavior data based on the display interface hot area distribution of the user terminal, the frequency intensity classification and the effectiveness label of the task progress state to generate a corresponding operation action sentence sequence; constructing a space-time alignment model of the text sentence sequence and the operation action sentence sequence, performing timestamp dynamic matching, game object association mapping and logical consistency detection to generate an intermediate feature set comprising a semantic keyword weight distribution, a decision preference intensity indicator and an emotional tendency index; performing intent inference on the intermediate feature set based on a dynamically updated context state matrix to determine an intent probability distribution, to generate a user behavior feature vector corresponding to the intent probability distribution, the context state matrix comprising a current task stage target, user historical behavior statistical features and a scene object attribute relationship graph.

3. The method of Claim 1, wherein, The rendering and generation of multi-modal interaction content matching the second behavior parameter set comprises: invoke a voice synthesis engine to generate voice waveform data with emotional prosody characteristics based on a dialogue response pattern in the second set of behavioral parameters; extract a basic action primitive from a preset action library according to a task guidance intensity value in the second set of behavioral parameters, perform action sequence mixing optimization on the basic action primitive through a physical simulation engine, generate three-dimensional character action data matching the task information, and generate interface highlight special effects positively correlated with the guidance intensity value; generate a facial micro-expression fusion weight based on an emotional expression coefficient in the second set of behavioral parameters, and control a particle special effect engine to generate environmental feedback special effects matching the emotional dimension based on the facial micro-expression fusion weight; input the voice waveform data, the three-dimensional character action data, the interface highlight special effects, and the environmental feedback special effects into a multi-modal synchronization engine to generate multi-modal interaction content with spatio-temporal feature alignment.

4. The method of Claim 3, wherein, The inputting of the voice waveform data, the three-dimensional character action data, the interface highlight special effects, and the environmental feedback special effects into the multi-modal synchronization engine to generate multi-modal interaction content with spatio-temporal feature alignment includes: adding a first level timestamp mark to the voice waveform data according to the timing requirements of the dialogue response pattern; adding a second level timestamp mark to the three-dimensional character action data based on the physical simulation logic of the task guidance intensity; adding a third level timestamp mark to the environmental feedback special effects according to the dynamic change curve of the emotional expression coefficient; aligning the first level, second level, and third level timestamp marks through a dynamic offset compensation algorithm to generate a unified reference time axis; inputting the voice waveform data, the three-dimensional character action data, the interface highlight special effects, and the environmental feedback special effects into the multi-modal synchronization engine according to the unified reference time axis for spatio-temporal feature alignment processing.

5. The method of interacting with an AI character within a game platform of claim 1, wherein, The optimization of the initial interaction strategy set through the strategy conflict resolution model to generate dynamic interaction strategies containing strategy priority labels and parameter adjustment amplitude values includes: building strategy conflict detection rules to detect logical contradictions and execution feasibility conflicts among dialogue guidance strategies, task assistance strategies, and emotional feedback strategies in the initial interaction strategy set, and generating corresponding conflict types; according to the detected conflict types, calling corresponding optimization strategies from a preset strategy optimization scheme library to iteratively optimize the initial interaction strategy set until all conflicts are eliminated; sorting and integrating the initial interaction strategy set after optimization according to the strategy priority labels and parameter adjustment amplitude values to generate the dynamic interaction strategies.

6. The method of interacting with an AI character within a game platform of claim 1, wherein, The displaying of the multi-modal interaction content to the user terminal includes: deploying a pre-generated cache pool of the multi-modal interaction content to an edge computing node connected to the user terminal, and rendering potential interaction content in the edge computing node; obtaining the current network delay measurement value of the user terminal to determine the corresponding data transmission mode; adjusting the multi-modal interaction content according to the data transmission mode, and sending the adjusted multi-modal interaction content to the user terminal for display.

7. The method of interacting with an AI character within a game platform of claim 1, wherein, The AI character interaction method further includes: Acquire a current game mode, the game mode including an online mode and a single mode; In a case where the game mode is in the online mode, establish a network connection relationship between a plurality of user terminals, and render multi-modal interactive content displayed by a current user terminal from a first perspective of the current user terminal, and synchronize game state data to all user terminals participating in online in real time; In a case where the game mode is in the single mode, disconnect the network connection with other user terminals, and run game logic on a local server of the user terminal, and store corresponding game progress to a local storage medium.

8. An interactive device, characterized by The interactive device includes a memory, a processor, and an AI character interaction program stored in the memory and executable on the processor within a game platform, the AI character interaction program within the game platform being configured to implement the AI character interaction method within the game platform according to any one of claims 1 to 7.

9. An interactive device, characterized by The interactive device according to claim 8. The interactive device according to claim 8.

Citation Information

Patent Citations

  • Intelligent NPC system interacting with players

    CN118807208A