Response interaction processing method and system based on artificial intelligence

Through the tensor neural tree model and fragmented falcon optimization algorithm, the problems of insufficient multimodal data fusion and interactive adaptability in existing technologies are solved, efficient semantic understanding and response generation are achieved, and the context understanding and response adaptability of artificial intelligence interactive systems are improved.

CN120686974AInactive Publication Date: 2025-09-23HUAIAN COLLEGE OF INFORMATION TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510772256.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing artificial intelligence interaction systems have bottlenecks in multimodal data fusion, structural modeling, interactive adaptability and response output, making it difficult to achieve efficient semantic understanding, accurate intent recognition and adapt to complex terminal environments.

Method used

The tensor neural tree model is combined with the fragmented falcon optimization algorithm, and through multimodal semantic feature modeling, time gating mechanism and asynchronous optimization strategy, semantic recognition and dynamic response generation in multi-round interaction scenarios are achieved.

Benefits of technology

It improves the accuracy and robustness of user intent recognition, enhances the system's response consistency and context adaptability in multi-round dialogue scenarios, and improves response generation efficiency and terminal adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120686974A_ABST
    Figure CN120686974A_ABST
Patent Text Reader

Abstract

The invention discloses a response interaction processing method and system based on artificial intelligence, and the method comprises the following steps: S1, collecting voice, image and text data inputted by a user, and generating multi-modal semantic feature representation; s2, inputting the multi-modal semantic feature representation into a tensor neural tree model, executing structure traversal and tensor fusion, introducing a time gating mechanism in combination with an interaction round and an intention state, and generating a context sensing representation; s3, generating response content based on the context sensing representation, and selecting adaptive channel output; s4, collecting user feedback data, generating an interaction log and extracting a behavior sequence and a response state; s5, constructing an optimization space of structural parameters, and optimizing the structural parameters by adopting a fragmented tenon hunting optimization algorithm; s6, the optimized tensor neural tree model is used for executing the steps from S1 to S3. According to the method, multi-round semantic understanding enhancement and response precision improvement are realized, and the method is suitable for an intelligent interaction system in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based response interaction processing method and system. Background Art

[0002] In current AI-driven interactive systems, multimodal data fusion and semantic understanding technologies are widely used in scenarios such as intelligent customer service, virtual assistants, and human-computer interaction terminals. Existing technologies primarily rely on traditional neural network structures or semantic recognition models based on the Transformer architecture to process user input, such as voice, images, and text, to identify intent and generate responses. However, these approaches generally suffer from rigid structures, limited context modeling capabilities, and insufficient interactive adaptability, making them difficult to support real-time intelligent responses and multi-round interactions in complex contexts.

[0003] When it comes to multimodal feature modeling, existing systems mostly combine speech, image, and text data in parallel or serial fashion. This lacks a unified and scalable structural fusion mechanism, which can easily lead to semantic redundancy or information weight imbalance, compromising overall feature extraction quality. Furthermore, the dynamic contextual changes in multimodal information across different interaction rounds are difficult to effectively model in existing models, often resulting in biased recognition of the user's true intent.

[0004] In terms of structural design, traditional deep learning models rely heavily on fixed hierarchies and static parameter control, limiting their ability to model structural dialogues. This is particularly true in multi-round semantic interactions, where they cannot fully capture the historical evolution of intent and semantic transfer relationships. Existing methods rarely consider information regulation mechanisms within model nodes and lack dynamic control strategies tailored to context and interaction state. This leads to unstable performance and limited generalization when faced with complex tasks or boundary conditions.

[0005] Most existing optimization algorithms for structural parameter optimization use general global optimization methods such as standard particle swarm optimization and genetic algorithms. While these methods offer some search capabilities, they suffer from limitations such as low search accuracy and a tendency to fall into local optima in high-dimensional structural parameter spaces. They often lack fine-grained control over local subspace heterogeneity, making it difficult to implement asynchronous adjustments and jump-like structural perturbations. This limits the adaptability and responsiveness of neural architectures in dynamic interactive tasks.

[0006] Furthermore, existing interactive processing systems lack adaptability in how they output response content. Different terminal environments (such as mobile devices, embedded systems, and low-bandwidth scenarios) require different response formats. However, most systems still use a unified channel or static output strategy, lacking dynamic channel selection mechanisms based on environmental conditions. This can lead to response delays and a reduced interactive experience.

[0007] In summary, existing response interaction processing methods all face varying degrees of technical bottlenecks in multimodal semantic fusion, structural modeling, interaction adaptation, and structural parameter optimization. To address these issues, a response interaction processing method and system with a unified multimodal fusion structure, dynamic context modeling capabilities, support for asynchronous optimization control strategies, and adaptability to complex terminal environments is urgently needed. This approach can improve the semantic understanding depth, response accuracy, and operational efficiency of the interactive system.

[0008] Therefore, how to provide an artificial intelligence-based response interaction processing method and system is an urgent problem that those skilled in the art need to solve. Summary of the Invention

[0009] One purpose of the present invention is to propose a response interaction processing method and system based on artificial intelligence. The present invention integrates multimodal semantic feature modeling, tensor neural tree structure expression and fragmented falcon optimization algorithm, and describes in detail the whole process of realizing semantic recognition, intent modeling and dynamic response generation in multi-round interaction scenarios. It has the advantages of strong context understanding ability, high structural adaptability and excellent response generation efficiency.

[0010] According to an embodiment of the present invention, a response interaction processing method based on artificial intelligence includes the following steps:

[0011] S1. Collect the user's input voice data, image data and text data, perform preprocessing to generate multimodal semantic feature representation, and record the current interaction round value;

[0012] S2. Input the multimodal semantic feature representation into the tensor neural tree model, perform tree structure traversal and tensor fusion operations, extract tensor outputs, perform intent recognition based on the tensor outputs, and generate the current intent state. A time gating mechanism is introduced in each structural node that performs the tensor fusion operation. The tensor output is adjusted according to the number of interaction rounds and the current intent state to generate a context-aware representation.

[0013] S3. Generate response content based on the context-aware representation. The response content includes text, voice, or image format. Select an adaptation channel from the preset output channels based on the current user's terminal type and network status to output the response content.

[0014] S4. Collect user feedback data, generate interaction logs, extract interaction behaviors, response acceptance states, and behavior sequences from the interaction logs, and construct an optimization space for the structural parameters of the tensor neural tree model. The structural parameters include tensor dimension, number of node layers, and gating control parameters.

[0015] S5. Use the fragmented falcon optimization algorithm to perform a global search on the structural parameters, perform fragmented local search, asynchronous raid scheduling and jump update operations, output the optimized structural parameters, and update the structural parameters of the tensor neural tree model;

[0016] S6. Use the updated tensor neural tree model to execute the interactive processing steps included in S1 to S3, receive new user input and generate response content.

[0017] Optionally, the preprocessing includes: performing acoustic feature extraction and semantic encoding on speech data, performing region segmentation and convolution feature encoding on image data, and performing word segmentation, part-of-speech tagging and context embedding on text data.

[0018] Optionally, the S2 specifically includes:

[0019] S21. Perform tree structure traversal, take the multimodal semantic feature representation as input, assign it to different structure nodes of the tensor neural tree model, traverse all structure nodes from top to bottom in the order of the tree hierarchy, and mark the traversed node index as (l,i), where l represents the number of layers and i represents the i-th node in the layer;

[0020] S22, perform tensor fusion operation at each traversed structure node, set the input tensor of the current node to X l,i , the fusion weight tensor is W l,i , then the computing node output tensor is Z l,i =f(W l,i ·X l,i ), where · represents tensor multiplication and f(·) represents the activation function;

[0021] S23, based on the node output tensor Z l,i Perform intent recognition, perform tensor dimension compression, feature aggregation, and embedding transformation operations, extract local intent representation vectors, and use the local intent representation vectors as the semantic response of the node in the current interaction round;

[0022] S24. The local intent representation vectors generated by all structural nodes are uniformly processed. The vectors output by each node are sequentially collected in the order of node traversal to form a vector set. A corresponding weighting coefficient is assigned to each vector in the vector set. The vectors are weighted and superimposed in a predetermined order to obtain a state vector representing the global semantic intent in the current interaction round.

[0023] S25. Introduce a time gating mechanism in each structural node that performs tensor fusion operations, obtain the current interaction round value and the intention state vector, set a scalar weight parameter for interaction time adjustment, and a set of semantic weight parameters corresponding to the dimensions of the intention state vector; multiply the interaction round value by the scalar weight parameter, multiply each component of the intention state vector by the corresponding component in the semantic weight parameter, and then sum them, and then add the two results to obtain an intermediate calculation value, input the intermediate calculation value into the nonlinear function mapping, obtain a gating strength value between zero and one, and adjust the retention ratio of the tensor output of the current structural node in the subsequent path;

[0024] S26. According to the time gating mechanism introduced in each structural node, the gating strength value obtained by jointly calculating the number of interaction rounds and the intention state vector through the set time weight parameters and semantic weight parameters is applied to the output tensor generated by the node in the tensor fusion operation, and the components of the output tensor are scaled and adjusted according to the gating strength to obtain the context-aware representation of the node.

[0025] Optionally, the step of selecting an adaptation channel from the preset output channels includes: determining the availability and bandwidth conditions of the output channel based on the terminal type and real-time network status currently used by the user, and selecting a target channel that meets the transmission requirements from the text, voice and image output channels to output the response content.

[0026] Optionally, the step of constructing the optimization space of structural parameters includes: extracting the user's interactive behavior, response acceptance status and behavior sequence from the interaction log, mapping them into tensor dimension configuration, structural layer configuration and gating adjustment strategy respectively, and combining them to form the optimization space of the tensor neural tree model structural parameters.

[0027] Optionally, the S5 specifically includes:

[0028] S51. Performing initial global sampling in the optimization space of the structural parameters of the tensor neural tree model, constructing a parameter combination set including tensor dimension configuration, node layer number configuration, and gate control parameter configuration, and dividing the structural parameter optimization space into a plurality of non-overlapping fragmented sub-regions;

[0029] S52, independently performing a local structure parameter search in each fragmented sub-region, and screening high-fitness candidate solutions in the region by combining the local fitness index to form a set of local candidate solutions with regional structural characteristics;

[0030] S53, performing asynchronous raid scheduling operations on each fragmented sub-region, and dynamically adjusting the search frequency, search step size, and perturbation strategy according to the fitness change rate, search convergence state, and local perturbation sensitivity of each sub-region during the search process;

[0031] S54. Based on the structural differences and distribution distances between the local candidate solution sets of each region, a jump update operation is performed between the plurality of fragmented sub-regions to generate a cross-region structural perturbation solution, thereby expanding the search range of the optimization space of the structural parameters.

[0032] S55. The local candidate solutions generated by all fragmented sub-regions are fused and evaluated with the structural perturbation solutions generated by jump perturbations. Based on the calculation of global fitness, the optimal structural parameter combination is selected as the optimized structural parameters, and the structural parameters of the tensor neural tree model are updated.

[0033] Optionally, the asynchronous raid scheduling operation in step S53 includes: monitoring the rate of change of the fitness of the local candidate solution in real time during the structural parameter search process for each fragmented sub-region, determining the current search stage based on the search convergence state and the response to the disturbance during the search process, and dynamically adjusting the search frequency, search step size, and disturbance strategy according to the search stage; wherein:

[0034] When the fitness shows an increase or decrease fluctuation amplitude greater than the first preset value in several consecutive iteration cycles, and the structural parameter combination is constantly changing, it is determined to be in the initial search stage, the search cycle is shortened, the step interval is extended, and the disturbance amplitude is increased;

[0035] When the fitness shows an increase or decrease fluctuation amplitude less than the second preset value in several consecutive iteration cycles, and the structural parameter combination does not undergo abrupt changes, it is determined to have entered the convergence stage, the search cycle is extended, the step size interval is shortened, and the disturbance amplitude is reduced;

[0036] When the fitness does not change in a preset number of consecutive iterations and the perturbation operation does not cause the structural parameter combination to be updated, it is determined to be a stagnation stage, and the sub-region search process is suspended or switched to the structural perturbation guided jumping mode;

[0037] Through the above mechanism, the search period, step interval and disturbance amplitude of each fragmented sub-region are regulated separately to achieve asynchronous optimization scheduling of the search process.

[0038] According to an embodiment of the present invention, an artificial intelligence-based response interaction processing system includes the following modules:

[0039] Multimodal preprocessing module: used to collect voice data, image data and text data and generate multimodal semantic feature representation;

[0040] Tensor neural tree processing module: used to receive multimodal semantic feature representations and perform tree structure traversal and tensor fusion operations;

[0041] Intent recognition module: used to perform intent recognition based on tensor fusion output and generate the current intent state;

[0042] Time-gated regulation module: used to regulate the tensor output in the structure node according to the number of interaction rounds and the current intention state;

[0043] Response generation module: used to generate response content based on context-aware representation and output it through a preset output channel;

[0044] Interaction log modeling module: used to collect user feedback and construct the optimization space of tensor neural tree model structure parameters;

[0045] Fragmentation optimization module: used to execute the fragmentation falcon optimization algorithm and output the optimized structural parameters;

[0046] Structure update module: used to update the structural parameters of the tensor neural tree model.

[0047] The beneficial effects of the present invention are:

[0048] First, the present invention introduces a tensor neural tree model to perform structured processing on the multimodal semantic feature representation, overcoming the problems of information redundancy and weight imbalance in traditional neural network structures when processing the fusion of speech, image and text, realizing the unified representation of different modal data and contextual semantic enhancement, and improving the accuracy and robustness of user intent recognition.

[0049] Secondly, the present invention introduces a time gating mechanism in the tensor fusion structure node, and dynamically adjusts the tensor output in combination with the interaction round value and the current intention state, effectively modeling the historical context and semantic migration characteristics of the interaction process, and improving the system's response consistency and context adaptability in multi-round dialogue scenarios.

[0050] In addition, the present invention adopts the fragmented falcon optimization algorithm to perform global search and jump update on the structural parameters of the tensor neural tree model, and has asynchronous raid scheduling and structural perturbation capabilities, which significantly improves the optimization efficiency of the model structural parameters and the depth of solution space exploration, thereby enhancing the model's generalization ability and real-time response performance in complex interactive environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0052] Figure 1 This is an overall flow chart of an artificial intelligence-based response interaction processing method proposed by the present invention. DETAILED DESCRIPTION

[0053] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0054] refer to Figure 1 , a response interaction processing method based on artificial intelligence, comprising the following steps:

[0055] S1. Collect the user's input voice data, image data and text data, perform preprocessing to generate multimodal semantic feature representation, and record the current interaction round value;

[0056] S2. Input the multimodal semantic feature representation into the tensor neural tree model, perform tree structure traversal and tensor fusion operations, extract tensor outputs, perform intent recognition based on the tensor outputs, and generate the current intent state. A time gating mechanism is introduced in each structural node that performs the tensor fusion operation. The tensor output is adjusted according to the number of interaction rounds and the current intent state to generate a context-aware representation.

[0057] S3. Generate response content based on the context-aware representation. The response content includes text, voice, or image format. Select an adaptation channel from the preset output channels based on the current user's terminal type and network status to output the response content.

[0058] S4. Collect user feedback data, generate interaction logs, extract interaction behaviors, response acceptance states, and behavior sequences from the interaction logs, and construct an optimization space for the structural parameters of the tensor neural tree model. The structural parameters include tensor dimension, number of node layers, and gating control parameters.

[0059] S5. Use the fragmented falcon optimization algorithm to perform a global search on the structural parameters, perform fragmented local search, asynchronous raid scheduling and jump update operations, output the optimized structural parameters, and update the structural parameters of the tensor neural tree model;

[0060] S6. Use the updated tensor neural tree model to execute the interactive processing steps included in S1 to S3, receive new user input and generate response content.

[0061] The present invention constructs a multimodal semantic response interaction processing method based on a tensor neural tree model, effectively integrating voice, image and text information, and combining the number of interaction rounds and intention status to realize context modeling. It breaks through the problems that traditional semantic models cannot dynamically perceive the interaction context and the structure is not scalable, and significantly improves the accuracy and coherence of the interactive system in intent recognition and response generation in complex contexts.

[0062] In this embodiment, the preprocessing includes: performing acoustic feature extraction and semantic encoding on speech data, performing region segmentation and convolution feature encoding on image data, and performing word segmentation, part-of-speech tagging and context embedding on text data.

[0063] In this embodiment, S2 specifically includes:

[0064] S21. Perform tree structure traversal, take the multimodal semantic feature representation as input, assign it to different structure nodes of the tensor neural tree model, traverse all structure nodes from top to bottom in the order of the tree hierarchy, and mark the traversed node index as (l,i), where l represents the number of layers and i represents the i-th node in the layer;

[0065] S22, perform tensor fusion operation at each traversed structure node, set the input tensor of the current node to X l,i , the fusion weight tensor is W l,i , then the computing node output tensor is Z l,i =f(W l,i ·X l,i ), where · represents tensor multiplication and f(·) represents the activation function;

[0066] S23, based on the node output tensor Z l,i Perform intent recognition, perform tensor dimension compression, feature aggregation, and embedding transformation operations, extract local intent representation vectors, and use the local intent representation vectors as the semantic response of the node in the current interaction round;

[0067] S24. The local intent representation vectors generated by all structural nodes are uniformly processed. The vectors output by each node are sequentially collected in the order of node traversal to form a vector set. A corresponding weighting coefficient is assigned to each vector in the vector set. The vectors are weighted and superimposed in a predetermined order to obtain a state vector representing the global semantic intent in the current interaction round.

[0068] S25. Introduce a time gating mechanism in each structural node that performs tensor fusion operations, obtain the current interaction round value and the intention state vector, set a scalar weight parameter for interaction time adjustment, and a set of semantic weight parameters corresponding to the dimensions of the intention state vector; multiply the interaction round value by the scalar weight parameter, multiply each component of the intention state vector by the corresponding component in the semantic weight parameter, and then sum them, and then add the two results to obtain an intermediate calculation value, input the intermediate calculation value into the nonlinear function mapping, obtain a gating strength value between zero and one, and adjust the retention ratio of the tensor output of the current structural node in the subsequent path;

[0069] S26. According to the time gating mechanism introduced in each structural node, the gating strength value obtained by jointly calculating the number of interaction rounds and the intention state vector through the set time weight parameters and semantic weight parameters is applied to the output tensor generated by the node in the tensor fusion operation, and the components of the output tensor are scaled and adjusted according to the gating strength to obtain the context-aware representation of the node.

[0070] This paper proposes a refined mechanism that integrates structural traversal, tensor fusion, local intent recognition, state vector construction, and time-gated regulation. By introducing hierarchical traversal and weight control strategies into each structural node, it not only achieves the orderly integration of multimodal semantic features within the structural hierarchy, but also jointly regulates interaction rounds and semantic states through a time-gated mechanism, effectively enhancing contextual perception and historical memory capabilities. This addresses the problems of existing semantic modeling methods, which lack dynamic regulation and cannot reflect the temporal evolution of interactions. It possesses the innovative advantages of clear structural hierarchy, strong response adaptability, and high semantic expression accuracy.

[0071] In this embodiment, the step of selecting an adaptation channel from the preset output channels includes: determining the availability and bandwidth conditions of the output channel based on the terminal type and real-time network status currently used by the user, and selecting a target channel that meets the transmission requirements from the text, voice, and image output channels to output the response content.

[0072] The present invention introduces an adaptation mechanism for preset output channels in the response generation stage, and dynamically selects text, voice or image response channels according to the user terminal type and network status. This solves the problem of a single response output mode and incompatibility with multi-scenario applications in existing interactive systems, and significantly improves the practicality and response efficiency of the system across devices and heterogeneous environments.

[0073] In this embodiment, the step of constructing the optimization space of structural parameters includes: extracting the user's interactive behavior, response acceptance status and behavior sequence from the interaction log, mapping them into tensor dimension configuration, structural layer configuration and gating adjustment strategy respectively, and combining them to form the optimization space of the tensor neural tree model structural parameters.

[0074] The present invention collects user feedback behaviors and constructs interaction logs, extracts multi-dimensional features such as behavior sequences and response acceptance states, effectively establishes a structural parameter optimization space, and provides data support for subsequent structural adaptation. This makes up for the problem that existing methods cannot introduce user feedback into the structural parameter adjustment path, and enhances the system's learning closed-loop capability and evolvability.

[0075] In this embodiment, the S5 specifically includes:

[0076] S51. Performing initial global sampling in the optimization space of the structural parameters of the tensor neural tree model, constructing a parameter combination set including tensor dimension configuration, node layer number configuration, and gate control parameter configuration, and dividing the structural parameter optimization space into a plurality of non-overlapping fragmented sub-regions;

[0077] S52, independently performing a local structure parameter search in each fragmented sub-region, and screening high-fitness candidate solutions in the region by combining the local fitness index to form a set of local candidate solutions with regional structural characteristics;

[0078] S53, performing asynchronous raid scheduling operations on each fragmented sub-region, and dynamically adjusting the search frequency, search step size, and perturbation strategy according to the fitness change rate, search convergence state, and local perturbation sensitivity of each sub-region during the search process;

[0079] S54. Based on the structural differences and distribution distances between the local candidate solution sets of each region, a jump update operation is performed between the plurality of fragmented sub-regions to generate a cross-region structural perturbation solution, thereby expanding the search range of the optimization space of the structural parameters.

[0080] S55. The local candidate solutions generated by all fragmented sub-regions are fused and evaluated with the structural perturbation solutions generated by jump perturbations. Based on the calculation of global fitness, the optimal structural parameter combination is selected as the optimized structural parameters, and the structural parameters of the tensor neural tree model are updated.

[0081] The present invention proposes a fragmented falcon optimization algorithm, which integrates fragmented local search, asynchronous raid scheduling and jump update mechanism to globally optimize the structural parameters of the tensor neural tree model. It solves the problem that traditional optimization algorithms converge slowly and fall into local optimality in high-dimensional structural parameter space, and has the advantages of high search efficiency, strong adaptability and good optimization accuracy.

[0082] In this embodiment, the asynchronous raid scheduling operation of step S53 includes: monitoring the rate of change of the fitness of the local candidate solution in real time during the structural parameter search process for each fragmented sub-region, determining the current search stage based on the search convergence state and the response to the disturbance during the search process, and dynamically adjusting the search frequency, search step size, and disturbance strategy according to the search stage; wherein:

[0083] When the fitness shows an increase or decrease fluctuation amplitude greater than the first preset value in several consecutive iteration cycles, and the structural parameter combination is constantly changing, it is determined to be in the initial search stage, the search cycle is shortened, the step interval is extended, and the disturbance amplitude is increased;

[0084] When the fitness shows an increase or decrease fluctuation amplitude less than the second preset value in several consecutive iteration cycles, and the structural parameter combination does not undergo abrupt changes, it is determined to have entered the convergence stage, the search cycle is extended, the step size interval is shortened, and the disturbance amplitude is reduced;

[0085] When the fitness does not change in a preset number of consecutive iterations and the perturbation operation does not cause the structural parameter combination to be updated, it is determined to be a stagnation stage, and the sub-region search process is suspended or switched to the structural perturbation guided jumping mode;

[0086] Through the above mechanism, the search period, step interval and disturbance amplitude of each fragmented sub-region are regulated separately to achieve asynchronous optimization scheduling of the search process.

[0087] The present invention introduces an asynchronous raid scheduling mechanism in the fragmented sub-regions, dynamically adjusts the search strategy based on the fitness change trend, convergence state and disturbance response, and realizes asynchronous resource allocation for different sub-regions. It breaks through the limitation of the unified scheduling strategy in the existing structural optimization process that is not suitable for personalized regions, and improves the adaptability and flexibility of the optimization strategy.

[0088] According to an embodiment of the present invention, an artificial intelligence-based response interaction processing system includes the following modules:

[0089] Multimodal preprocessing module: used to collect voice data, image data and text data and generate multimodal semantic feature representation;

[0090] Tensor neural tree processing module: used to receive multimodal semantic feature representations and perform tree structure traversal and tensor fusion operations;

[0091] Intent recognition module: used to perform intent recognition based on tensor fusion output and generate the current intent state;

[0092] Time-gated regulation module: used to regulate the tensor output in the structure node according to the number of interaction rounds and the current intention state;

[0093] Response generation module: used to generate response content based on context-aware representation and output it through a preset output channel;

[0094] Interaction log modeling module: used to collect user feedback and construct the optimization space of tensor neural tree model structure parameters;

[0095] Fragmentation optimization module: used to execute the fragmentation falcon optimization algorithm and output the optimized structural parameters;

[0096] Structure update module: used to update the structural parameters of the tensor neural tree model.

[0097] The system modules described in the present invention are clearly divided, covering the entire process of multimodal semantic feature generation, tensor structure processing, intent recognition, response output and structural parameter optimization. The module functions are consistent with the data flow, and a response interaction system with high scalability and high integration is constructed, which solves the problems of loose structure and serious functional coupling of existing systems.

[0098] Example 1:

[0099] In order to verify the feasibility of the present invention in implementation, the present invention is applied to a large-scale intelligent customer service platform, aiming to improve the platform's semantic understanding ability, response consistency and service adaptation efficiency in multiple rounds of voice-image-text mixed interactions. The platform processes more than 150,000 user requests per day, and users cover a variety of terminal types, including mobile APPs, web pages, embedded devices and intelligent voice terminals. When faced with multiple rounds of continuous questions, ambiguous language input and mixed image and text requests, existing systems often fail to respond to user intentions due to insufficient semantic modeling accuracy, broken context understanding or improper response adaptation, resulting in overall low customer satisfaction.

[0100] The platform's existing technology, based on a standard Transformer model and a parallel multimodal fusion structure, achieved an intent recognition accuracy of approximately 71.2% for mixed text and image requests. Its ability to maintain consistency across multiple conversation rounds was poor, with accuracy dropping dramatically in conversation rounds beyond the fourth. To address this issue, the present invention introduces a tensor neural tree model structure, which implements hierarchical fusion of multimodal semantic features within structural nodes through tree-like structure traversal. A time-gating mechanism is introduced within the fusion nodes, which uses the interaction round count and the current semantic state vector to jointly regulate tensor output. This improves the model's ability to perceive contextual coherence and discern the user's true intent.

[0101] During actual deployment, user input is first preprocessed by the speech recognition module, image segmentation module, and text semantic encoding module to generate a unified multimodal semantic feature representation, which is then fed into a tensor neural tree model. The system automatically assigns the input features to structural nodes and performs tensor fusion. During the fusion process, the output of each structural node performs intent recognition, outputting a local semantic response. The global weighted aggregation module then generates the current interaction state vector. This state vector is then used together with the interaction round count to calculate the gate strength, dynamically adjusting the tensor participation in the output path, thereby achieving context-aware structural output regulation.

[0102] During the model operation phase, the system monitors multi-dimensional indicators such as response accuracy, intent recognition accuracy, and response channel adaptation rate in different interaction rounds. Compared with the traditional model, the system of the present invention maintains an intent retention rate of more than 89% in the 1st to 6th rounds of dialogue, of which the accuracy rate in the 4th round is still as high as 87.1%, significantly higher than the 64.3% of the original system. At the same time, in the case of incomplete voice or fuzzy image input, the system can restore the user's intention with the help of the context state vector and generate logically coherent mixed text and image response content. In addition, the output channel selection mechanism combines the terminal network status to dynamically select image output or text output, which increases the terminal adaptation accuracy to 98.5%.

[0103] To verify the system's generalization capabilities, the platform was deployed and data collected in 15 business scenarios (such as user consultation, order changes, after-sales applications, etc.). After a 14-day operation cycle, a customer satisfaction survey showed that the average score of the business scenarios using the system of the present invention reached 4.71 points (out of 5 points), an increase of 18.4% over the original system. At the same time, the average response delay was controlled within 1.23 seconds, and the average system load increased stability by about 21%, demonstrating high practical deployment feasibility and algorithm stability.

[0104] The following is some real data collected during the experiment, covering the comparison of system operation indicators.

[0105] Table 1: Summary of comparative experimental results between the proposed model and the original system on typical interaction indicators

[0106]

[0107]

[0108] According to the table "Summary of Comparative Experimental Results of the Model of the Present Invention and the Original System on Typical Interaction Indicators", it can be seen that the system of the present invention has achieved significant improvements in multiple key performance indicators compared to the original system. Figure 1 In terms of consistency, especially in the fourth round of continuous dialogue, the model of the present invention maintained an accuracy of 87.1%, which is significantly higher than the 64.3% of the original system, indicating that it has stronger capabilities in long-term context retention and semantic inheritance.

[0109] Secondly, the recognition accuracy of mixed text and image requests and the accuracy of fuzzy input responses were improved by 18.6% and 19.5% respectively. This shows that the tensor neural tree structure and its fusion mechanism of the present invention can more effectively process unstructured input and semantically incomplete information, and have strong robustness and fault tolerance. In addition, the output channel adaptation accuracy was increased from 82.6% of the original system to 98.5%, proving that the introduced response channel selection mechanism shows extremely high adaptability when the terminal type and network status change, and effectively improves the problem of response content output deviation.

[0110] In terms of system efficiency, this invention reduces average response time from 1.67 seconds to 1.23 seconds, optimizing overall interaction latency and improving response fluency. Furthermore, the system's peak stability improves by over 21%, meaning the model maintains structural stability and task processing capabilities even in high-concurrency environments.

[0111] Finally, from a user perspective, customer satisfaction scores increased from 3.98 to 4.71, directly demonstrating their high recognition of the system's semantic understanding, interactive coherence, and service quality. In summary, the tabular data fully validates the technical innovation and engineering value of this invention in terms of structural modeling, optimization algorithms, and system output.

[0112] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A response interaction processing method based on artificial intelligence, characterized in that: The steps include: S1. Collect the user's input voice data, image data and text data, perform preprocessing to generate multimodal semantic feature representation, and record the current interaction round value; S2. Input the multimodal semantic feature representation into the tensor neural tree model, perform tree structure traversal and tensor fusion operations, extract tensor outputs, perform intent recognition based on the tensor outputs, and generate the current intent state. A time gating mechanism is introduced in each structural node that performs the tensor fusion operation. The tensor output is adjusted according to the number of interaction rounds and the current intent state to generate a context-aware representation. S3. Generate response content based on the context-aware representation. The response content includes text, voice, or image format. Select an adaptation channel from the preset output channels based on the current user's terminal type and network status to output the response content. S4. Collect user feedback data, generate interaction logs, extract interaction behaviors, response acceptance states, and behavior sequences from the interaction logs, and construct an optimization space for the structural parameters of the tensor neural tree model. The structural parameters include tensor dimension, number of node layers, and gating control parameters. S5. Use the fragmented falcon optimization algorithm to perform a global search on the structural parameters, perform fragmented local search, asynchronous raid scheduling and jump update operations, output the optimized structural parameters, and update the structural parameters of the tensor neural tree model; S6. Use the updated tensor neural tree model to execute the interactive processing steps included in S1 to S3, receive new user input and generate response content.

2. A response interaction processing method based on artificial intelligence according to claim 1, characterized in that: The preprocessing includes: performing acoustic feature extraction and semantic encoding on speech data, performing region segmentation and convolution feature encoding on image data, and performing word segmentation, part-of-speech tagging and context embedding on text data.

3. A response interaction processing method based on artificial intelligence according to claim 2, characterized in that: The S2 specifically includes: S21. Perform tree structure traversal, take the multimodal semantic feature representation as input, assign it to different structure nodes of the tensor neural tree model, traverse all structure nodes from top to bottom in the order of the tree hierarchy, and mark the traversed node index as (l,i), where l represents the number of layers and i represents the i-th node in the layer; S22, perform tensor fusion operation at each traversed structure node, set the input tensor of the current node to X l,i , the fusion weight tensor is W l,i , then the computing node output tensor is Z l,i =f(W l,i ·X l,i ), where · represents tensor multiplication and f(·) represents the activation function; S23, based on the node output tensor Z l,i Perform intent recognition, perform tensor dimension compression, feature aggregation, and embedding transformation operations, extract local intent representation vectors, and use the local intent representation vectors as the semantic response of the node in the current interaction round; S24. The local intent representation vectors generated by all structural nodes are uniformly processed. The vectors output by each node are sequentially collected in the order of node traversal to form a vector set. A corresponding weighting coefficient is assigned to each vector in the vector set. The vectors are weighted and superimposed in a predetermined order to obtain a state vector representing the global semantic intent in the current interaction round. S25. Introduce a time gating mechanism in each structural node that performs tensor fusion operations, obtain the current interaction round value and the intention state vector, set a scalar weight parameter for interaction time adjustment, and a set of semantic weight parameters corresponding to the dimensions of the intention state vector; multiply the interaction round value by the scalar weight parameter, multiply each component of the intention state vector by the corresponding component in the semantic weight parameter, and then sum them, and then add the two results to obtain an intermediate calculation value, input the intermediate calculation value into the nonlinear function mapping, obtain a gating strength value between zero and one, and adjust the retention ratio of the tensor output of the current structural node in the subsequent path; S26. According to the time gating mechanism introduced in each structural node, the gating strength value obtained by jointly calculating the number of interaction rounds and the intention state vector through the set time weight parameters and semantic weight parameters is applied to the output tensor generated by the node in the tensor fusion operation, and the components of the output tensor are scaled and adjusted according to the gating strength to obtain the context-aware representation of the node.

4. A response interaction processing method based on artificial intelligence according to claim 3, characterized in that: The step of selecting an adaptation channel from the preset output channels includes: determining the availability and bandwidth conditions of the output channel based on the terminal type currently used by the user and the real-time network status, and selecting a target channel that meets the transmission requirements from the text, voice and image output channels to output the response content.

5. A response interaction processing method based on artificial intelligence according to claim 4, characterized in that: The step of constructing the optimization space of structural parameters includes: extracting the user's interactive behavior, response acceptance status and behavior sequence from the interaction log, mapping them into tensor dimension configuration, structural layer configuration and gating adjustment strategy respectively, and combining them to form the optimization space of the tensor neural tree model structural parameters.

6. A response interaction processing method based on artificial intelligence according to claim 5, characterized in that: The S5 specifically includes: S51. Performing initial global sampling in the optimization space of the structural parameters of the tensor neural tree model, constructing a parameter combination set including tensor dimension configuration, node layer number configuration, and gate control parameter configuration, and dividing the structural parameter optimization space into a plurality of non-overlapping fragmented sub-regions; S52, independently performing a local structure parameter search in each fragmented sub-region, and screening high-fitness candidate solutions in the region by combining the local fitness index to form a set of local candidate solutions with regional structural characteristics; S53, performing asynchronous raid scheduling operations on each fragmented sub-region, and dynamically adjusting the search frequency, search step size, and perturbation strategy according to the fitness change rate, search convergence state, and local perturbation sensitivity of each sub-region during the search process; S54. Based on the structural differences and distribution distances between the local candidate solution sets of each region, a jump update operation is performed between the plurality of fragmented sub-regions to generate a cross-region structural perturbation solution, thereby expanding the search range of the optimization space of the structural parameters. S55. The local candidate solutions generated by all fragmented sub-regions are fused and evaluated with the structural perturbation solutions generated by jump perturbations. Based on the calculation of global fitness, the optimal structural parameter combination is selected as the optimized structural parameters, and the structural parameters of the tensor neural tree model are updated.

7. A response interaction processing method based on artificial intelligence according to claim 6, characterized in that: The asynchronous raid scheduling operation in step S53 includes: monitoring the rate of change of the fitness of the local candidate solution in real time during the structural parameter search process for each fragmented sub-region, determining the current search stage based on the search convergence state and the response to the disturbance during the search process, and dynamically adjusting the search frequency, search step size, and disturbance strategy according to the search stage; wherein: When the fitness shows an increase or decrease fluctuation amplitude greater than the first preset value in several consecutive iteration cycles, and the structural parameter combination is constantly changing, it is determined to be in the initial search stage, the search cycle is shortened, the step interval is extended, and the disturbance amplitude is increased; When the fitness shows an increase or decrease fluctuation amplitude less than the second preset value in several consecutive iteration cycles, and the structural parameter combination does not undergo abrupt changes, it is determined to have entered the convergence stage, the search cycle is extended, the step size interval is shortened, and the disturbance amplitude is reduced; When the fitness does not change in a preset number of consecutive iterations and the perturbation operation does not cause the structural parameter combination to be updated, it is determined to be a stagnation stage, and the sub-region search process is suspended or switched to the structural perturbation guided jumping mode; Through the above mechanism, the search period, step interval and disturbance amplitude of each fragmented sub-region are regulated separately to achieve asynchronous optimization scheduling of the search process.

8. An artificial intelligence-based response interaction processing system, applied to an artificial intelligence-based response interaction processing method according to any one of claims 1 to 7, characterized in that: Includes the following modules: Multimodal preprocessing module: used to collect voice data, image data and text data and generate multimodal semantic feature representation; Tensor neural tree processing module: used to receive multimodal semantic feature representations and perform tree structure traversal and tensor fusion operations; Intent recognition module: used to perform intent recognition based on tensor fusion output and generate the current intent state; Time-gated regulation module: used to regulate the tensor output in the structure node according to the number of interaction rounds and the current intention state; Response generation module: used to generate response content based on context-aware representation and output it through a preset output channel; Interaction log modeling module: used to collect user feedback and construct the optimization space of tensor neural tree model structure parameters; Fragmentation optimization module: used to execute the fragmentation falcon optimization algorithm and output the optimized structural parameters; Structure update module: used to update the structural parameters of the tensor neural tree model.

Citation Information

Cited By

  • Spatial intelligent multi-modal tree planning method and system based on value-driven cutting

    CN121477901A

  • Order process merging workshop scheduling method and system based on job level constraint

    CN121961098A