Object service assistance method and apparatus based on multi-modal interaction, device, and medium

CN122765079APending Publication Date: 2026-09-15CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610910096.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-15

Smart Images

  • Figure CN122765079A_ABST
    Figure CN122765079A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent decision-making, and discloses an object service assisting method and device based on multi-modal interaction, equipment and a medium, which comprises the following steps: extracting structured perception feature data through an end-side multi-modal model; calling knowledge resource data and object service basic data through a server-side control model, generating service intention data, an image and service decision data; determining a target intelligent agent set based on the service decision data and generating assisting action data; collecting response feedback data, updating the image and generating subsequent service decision data. The application can be applied to business scenes such as financial technology and medical health, and through continuous processing of end-side perception, server-side decision-making, intelligent agent action and feedback updating, the service decision is updated with feedback changes, and the continuity and dynamic adjustment capability of the object service assisting process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, and in particular to an object service assistance method, apparatus, device and medium based on multimodal interaction. Background Technology

[0002] Existing service support tools mostly remain at the level of passive querying and static display. Service personnel need to manually obtain object information, business notifications, resource status, and historical interaction records from multiple data sources, and then make service judgments based on experience. Because the current interaction content, object status, service personnel status, and business knowledge resources are difficult to connect continuously in the same processing chain, existing tools cannot make timely service decisions that match the current scenario, and subsequent service actions are also difficult to continuously update based on feedback, resulting in judgment lags and action breakpoints in the service support process.

[0003] In the fintech sector, businesses such as banking, insurance, securities, and wealth management typically involve financial customer profiling, product benefits, compliance policies, market dynamics, risk warnings, and service personnel task assignments. Existing financial sales tools often rely on service personnel manually querying financial customer data, filtering financial product information, and searching for policy notices. This makes it difficult to link the current interaction content of financial customers, the current status of financial customer service, and the current status of service personnel for service decisions, resulting in a disconnect between financial customer service actions and real-time business scenarios.

[0004] In the healthcare sector, services such as health management, chronic disease follow-up, physical examination management, medication reminders, and health consultations typically involve health records, follow-up records, examination reports, health risk alerts, service resources, and the task status of healthcare personnel. Existing healthcare service tools often focus on fixed-process reminders or single data queries, making it difficult to combine user interaction content, changes in health status, and the current status of healthcare personnel to form continuous service judgments. This results in health reminders, follow-up arrangements, and service resource allocation being difficult to adjust in a timely manner based on feedback. Summary of the Invention

[0005] The main objective of this invention is to provide an object service assistance method, apparatus, device, and storage medium based on multimodal interaction, aiming to solve the technical problem that existing service assistance tools are unable to form a continuous processing chain from scene perception, service decision-making, action generation to feedback updates, resulting in service decisions and service actions being difficult to dynamically adjust with the actual service process.

[0006] To achieve the above objectives, the present invention provides an object service assistance method based on multimodal interaction, comprising: Acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the terminal-side multimodal model; The structured perception feature data is sent to the server-side control model, and the knowledge resource data and object service basic data are retrieved through the server-side control model. Through the server-side control model, service intent data, target object service profiles, and service personnel status profiles are generated based on the structured perception feature data, the knowledge resource data, and the object service basic data. Service decision data is generated based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. Based on the service decision data, a set of target intelligent agents is determined, auxiliary action data is generated through the set of target intelligent agents, and the auxiliary action data is sent to the interactive terminal; Collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

[0007] Furthermore, to achieve the above objectives, the present invention provides an object service auxiliary device based on multimodal interaction, comprising: The edge-side multimodal perception module is used to acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the edge-side multimodal model; The server-side resource retrieval module is used to send the structured perception feature data to the server-side control model, and retrieve knowledge resource data and object service basic data through the server-side control model; The intent profile generation module is used to generate service intent data, target object service profiles, and service personnel status profiles based on the structured perception feature data, the knowledge resource data, and the object service basic data through the server-side control model. The service decision generation module is used to generate service decision data based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. The intelligent agent action orchestration module is used to determine a target intelligent agent set based on the service decision data, generate auxiliary action data through the target intelligent agent set, and send the auxiliary action data to the interactive terminal; The feedback update module is used to collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multimodal interaction-based object service auxiliary program stored in the memory and executable on the processor, wherein when the multimodal interaction-based object service auxiliary program is executed by the processor, it implements the steps of the multimodal interaction-based object service auxiliary method as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an object service assistance program based on multimodal interaction, wherein the object service assistance program based on multimodal interaction, when executed by a processor, implements the steps of the object service assistance method based on multimodal interaction as described above.

[0010] Beneficial Effects: This invention relates to the field of intelligent decision-making technology, and discloses a method, apparatus, device, and medium for assisting object services based on multimodal interaction. The method includes: acquiring multimodal interaction input data collected by an interactive terminal; extracting structured perception feature data through an edge-side multimodal model; sending the structured perception feature data to a server-side control model and retrieving knowledge resource data and object service basic data; generating service intent data, target object service profiles, and service personnel status profiles, and generating service decision data accordingly; determining a set of target intelligent agents based on the service decision data, and generating auxiliary action data through the target intelligent agent set; collecting response feedback data, updating the target object service profile and service personnel status profile, and generating subsequent service decision data. This invention can be applied to business scenarios such as fintech and healthcare. Through continuous processing of edge-side perception, server-side decision-making, intelligent agent action, and feedback updates, service decision data can be updated along with response feedback data, reducing service lag caused by manual retrieval and experience-based judgment, and improving the continuity and dynamic adjustment capability of the object service assistance process. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an object service assistance method based on multimodal interaction according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the object service assistance method based on multimodal interaction according to the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the object service auxiliary device based on multimodal interaction of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The object service assistance method based on multimodal interaction provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain multimodal interactive input data collected by the client's interactive terminal, extract structured perception feature data through the edge-side multimodal model, send the structured perception feature data to the server-side control model, and retrieve knowledge resource data and object service basic data; generate service intent data, target object service profiles, and service personnel status profiles, and generate service decision data accordingly; determine the target intelligent agent set based on the service decision data, and generate auxiliary action data through the target intelligent agent set; collect response feedback data, update the target object service profile and service personnel status profile, and generate subsequent service decision data. This invention can be applied to business scenarios such as fintech and healthcare. Through continuous processing of edge-side perception, server-side decision-making, intelligent agent action, and feedback updates, service decision data can be updated with response feedback data, reducing service lag caused by manual retrieval and experience-based judgment, and improving the continuity and dynamic adjustment capability of the object service auxiliary process. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention will now be described in detail through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the object service assistance method based on multimodal interaction provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the object service assistance method based on multimodal interaction proposed in this invention includes the following steps: S10, acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the terminal-side multimodal model; In this embodiment, the interactive terminal collects input content on the interface side where service personnel interact with the target object. The collected content is not limited to single text input, but also includes input content that reflects the current service scenario, such as voice, images, page operations, button triggers, pause behaviors, withdrawal behaviors, and confirmation behaviors. Multimodal interactive input data is formed by aggregating input content from different modalities according to the same interaction round. During the aggregation process, collection time, modality source, and session markers are configured for different input content, enabling the edge-side multimodal model to determine whether different input content belongs to the same service context. Voice content is encoded into semantic segments by edge-side voice encoding, text content is encoded into semantic segments by edge-side text encoding, image content is encoded into object segments and text segments by edge-side visual encoding, and operation behaviors are encoded into behavior segments by behavior sequence encoding. The edge-side multimodal model aligns the above segments according to the collection time and session markers, and merges the content from voice, text, images, and behaviors that points to the same target object or the same service item to form a fielded result that can express service intent candidates, object pointing, personnel status prompts, scenario trigger prompts, and collection status. Structured perceptual feature data consists of field-based results, which may include conversation fields, modality fields, time fields, intent fields, object fields, personnel status fields, scene trigger fields, and acquisition status fields. When the edge multimodal model adopts a structure of multiple encoders plus a fusion layer, the speech encoder, text encoder, visual encoder, and behavior encoder each output modal representations. The fusion layer merges the modal representations based on time intervals, object references, and semantic similarity. The output layer then converts the merged result into structured perceptual feature data.

[0016] Under conditions of sufficient computing power and stable network conditions on the interactive terminal, an edge-side multimodal model is employed, in which speech encoders, text encoders, visual encoders, and behavior encoders operate in parallel, to extract structured perceptual feature data. Financial service terminals collect voice input from financial customers regarding account services, wealth management products, insurance benefits, and risk warnings. Simultaneously, they collect text notes entered by service personnel, product page screenshots, and interface operation behaviors. The edge-side multimodal model integrates the consultation intent in the voice, the service items in the text, the product identifiers in the screenshots, and the page dwell information in the operation behaviors to form structured perceptual feature data.

[0017] In situations where the computing power of the interactive terminal is limited or the network is unstable, a lightweight edge-side multimodal model is adopted to reduce the edge-side load. Voice content is first converted into short text semantic fragments, and only the text area and object identifier area of ​​the image content are extracted. Operational behaviors are compressed and encoded according to page number, trigger type, and dwell time. The healthcare service terminal collects user voice or text content regarding medication reminders, health indicators, and follow-up matters, while simultaneously collecting images of medical examination reports and reminder confirmation actions. The edge-side multimodal model integrates health consultation intent, report indicator objects, and behavioral response status into structured perceptual feature data, avoiding the transmission of complete voice and images to the terminal side.

[0018] This embodiment collects multimodal interactive input data through an interactive terminal, and the terminal-side multimodal model completes the encoding, alignment, and field-based output of voice, text, images, and operational behaviors. Dispersed input content can be converted into structured perceptual feature data. Since the structured perceptual feature data already contains directly identifiable fields such as intent, object, personnel status, scene triggering, and collection status, subsequent processing can reduce repeated parsing of multi-source input content, reduce the lag caused by manual information processing, and improve the timeliness and consistency of service scene perception.

[0019] S20, the structured perception feature data is sent to the server control model, and the knowledge resource data and object service basic data are retrieved through the server control model; In this embodiment, the structured perception feature data is encapsulated into a data packet on the interactive terminal side. The data packet retains fields such as session marker, modality source, acquisition time, intent candidate, object pointer, personnel status prompt, scene trigger prompt, and acquisition status, and adds field integrity markers and data source markers. The interactive terminal sends the data packet through a remote interface, message queue, or end-to-cloud synchronization channel. After parsing the data packet, the server-side control model forms a resource retrieval index based on the object pointer, intent candidate, scene trigger prompt, and personnel status prompt.

[0020] Knowledge resource data includes scenario knowledge, organizational knowledge, service resources, and operational guidelines. In fintech business scenarios, this may include financial product descriptions, rights and benefits statements, compliance notices, and risk warnings. In healthcare business scenarios, it may include health service guidelines, chronic disease management tips, medication reminders, and follow-up process instructions. Object service basic data includes target object basic records, interaction records, service status records, service personnel capability records, and task status records. The server-side control model retrieves knowledge resource data and object service basic data separately based on resource retrieval indexes, ensuring that both types of data maintain the same session tagging with the structured perception feature data.

[0021] With a stable connection between the interactive terminal and the server, structured perception feature data is sent via a remote interface request. After financial service personnel collect customer consultation voice messages, product page screenshots, and operational behaviors through mobile terminals, the server-side control model retrieves financial product descriptions, benefit statements, compliance notices, and customer service records based on the target object and scenario-triggered prompts.

[0022] In situations where the network status of the interactive terminal is unstable, a combination of local caching and message queues is used to transmit structured perception feature data. After the health management terminal collects the user's health consultation text, report images, and reminder confirmation behavior, the server-side control model retrieves health service guidelines, follow-up records, service status records, and personnel task status records based on intent candidates and object pointers.

[0023] In the fintech business, the data generated by the account manager's terminal includes the financial customer target, the intent to inquire about rights and interests, and the source of the product page. Based on this, the server-side control model retrieves product rights and interests descriptions, risk warnings, handling instructions, and customer interaction records.

[0024] In the healthcare business field, the data generated by health management terminals includes user target information, medication reminder intent, and report image source. The server-side control model uses this information to retrieve medication reminder instructions, health service guidelines, follow-up records, and service personnel task status.

[0025] This embodiment retrieves knowledge resource data and object service basic data based on structured perception feature data through a server-side control model. It can obtain matching resource content and basic status data around the same interaction round, reduce the delay of manual retrieval, and improve the accuracy of resource retrieval and the continuity of service processing.

[0026] S30, through the server-side control model, service intent data, target object service profile, and service personnel status profile are generated based on the structured perception feature data, the knowledge resource data, and the object service basic data; In this embodiment, after receiving structured perception feature data, knowledge resource data, and object service basic data, the server-side control model aligns the fields according to session tags, object pointers, and scene trigger prompts. The structured perception feature data provides intent candidates, object pointers, personnel status prompts, and collection status; the knowledge resource data provides scene knowledge, service resources, and operation guidance; and the object service basic data provides target object records, interaction records, service status records, and service personnel status records.

[0027] The server-side control model generates service intent data by combining intent candidates, scenario trigger prompts, and knowledge resource data. It then generates a target object service profile by combining object targeting, object service basic data, and knowledge resource data. Finally, it generates a service personnel status profile by combining personnel status prompts, object service basic data, and knowledge resource data. Service intent data may include service category, service goal, and urgency level; the target object service profile may include object activity status, demand status, and service gaps; and the service personnel status profile may include capability status, task load, and response preferences.

[0028] In financial service scenarios, to improve the stability of product consultation and rights service assessment, a combination of classification models and vector retrieval is adopted. The server-side control model retrieves product descriptions, rights descriptions, risk warnings, and customer service records based on the financial customer target, consultation intent, and product page source, generating service intent data, target customer service profiles, and service personnel status profiles.

[0029] In health management scenarios, a combination of semantic reasoning models and profile field templates is used to handle health consultations and follow-ups. The server-side control model retrieves medication reminder instructions, health service guidelines, follow-up records, and personnel task status based on the user object, medication reminder intent, and report image source, generating service intent data, target object service profiles, and service personnel status profiles.

[0030] This embodiment integrates structured perception feature data, knowledge resource data, and object service basic data through a server-side control model. The edge perception results, knowledge content, and basic status can jointly participate in intent recognition and profile generation, reducing judgment bias caused by a single data source and improving the matching between service intent data, target object service profile, and service personnel status profile.

[0031] S40, through the server-side control model, service decision data is generated based on the service intent data, the target object service profile, and the service personnel status profile; In this embodiment, after receiving service intent data, target object service profiles, and service personnel status profiles, the server-side control model aligns the fields of these three types of data. The service intent data provides the service category, service goal, triggering scenario, and urgency level. The target object service profile provides the object's activity status, demand status, service gap, and resource preferences. The service personnel status profile provides capability status, task load, response preferences, and service availability. The server-side control model compares the service goal with the object's demand status, the service gap with the capability status, and the urgency level with the task load, forming a data combination that can be used for decision-making.

[0032] Service decision data is used to carry the decision results required for subsequent agent invocation and auxiliary action generation. The server-side control model can write service action candidate data, service priority data, agent invocation conditions, and auxiliary action output conditions into the service decision data. Service action candidate data represents the types of service actions that can be generated, service priority data represents the processing order of different service actions, agent invocation conditions represent the capability range of the agent to be invoked subsequently, and auxiliary action output conditions represent the display and triggering form of the auxiliary action on the interactive terminal.

[0033] In financial service scenarios, to generate service decision data that better reflects the customer's situation, a combination of profile field comparison and semantic classification models is used. The server-side control model generates service action candidate data and service priority data based on the financial customer's intent to inquire about rights and interests, product demand status, account service status, and account manager's task load, and generates intelligent agent invocation conditions for calling intelligent agents for product description, rights and interests reminder, or communication assistance.

[0034] In health management scenarios, to generate service decision data suitable for follow-up and reminders, a combination of intent-layered recognition and state threshold assessment is adopted. The server-side control model generates candidate data for service actions such as follow-up reminders, health service gaps, follow-up status, and service personnel workload based on the user's medication reminder intent, health service gaps, follow-up status, and service personnel workload. It also generates auxiliary action output conditions, enabling subsequent auxiliary actions to be presented in the form of reminder cards, interactive Q&As, or task prompts.

[0035] This embodiment integrates service intent data, target object service profiles, and service personnel status profiles through a server-side control model. Service targets, object needs, and personnel carrying status can jointly participate in the generation of service decision data, reducing the bias caused by making judgments based on a single intent or profile, and improving the accuracy of service action selection and agent invocation.

[0036] S50, determine a set of target intelligent agents based on the service decision data, generate auxiliary action data through the set of target intelligent agents, and send the auxiliary action data to the interactive terminal; In this embodiment, the service decision data includes service action candidate data, service priority data, agent invocation conditions, and auxiliary action output conditions. Based on the agent invocation conditions, the server-side control model filters agents capable of handling the current service action from the registered agent capability information and determines the target agent set by combining the service priority data. The target agent set may include one or more of the following types: task generation agents, resource extraction agents, and interactive prompt agents.

[0037] After receiving candidate service action data, the target intelligent agent set generates task action fragments, resource action fragments, or interactive prompt fragments. The server-side control model arranges the action fragments in order according to service priority data and organizes them into auxiliary action data based on auxiliary action output conditions. Auxiliary action data may include to-do reminders, resource recommendations, communication prompts, simulated Q&A content, or page-triggered content, and is sent to the interactive terminal in the form of terminal display data, terminal interaction trigger data, etc.

[0038] In situations requiring the simultaneous generation of service reminders and communication support content, a multi-agent collaborative approach is employed to generate auxiliary action data. In financial service scenarios, the server-side control model determines task-generating agents and interaction-prompt agents based on service decision data. The task-generating agents generate customer follow-up tasks, while the interaction-prompt agents generate rights explanation scripts and risk warning content. The server-side control model then merges these two types of results into auxiliary action data and sends it to the account manager's terminal.

[0039] In scenarios where only a single prompt is needed, a single agent invocation approach is used to reduce response latency. In health management scenarios, the server-side control model determines the interactive prompt agent based on service decision data. The interactive prompt agent generates medication reminders or follow-up prompts, and the server-side control model converts the generated results into reminder cards and sends them to the health management terminal.

[0040] This embodiment determines the target intelligent agent set through service decision data, and generates auxiliary action data from the target intelligent agent set. Service actions can be allocated according to the intelligent agent's capabilities and service priorities, reducing the time spent on manually selecting tools and organizing prompts, and enabling the interactive terminal to receive action prompts that are more in line with the current service scenario.

[0041] S60, collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

[0042] In this embodiment, after the auxiliary action data is sent to the interactive terminal, the interactive terminal collects the service personnel's operations, the target object's response, and the terminal's interaction results to form response feedback data. The response feedback data may include viewing status, confirmation status, reply content, task completion status, page dwell time status, and reminder closed status. The server-side control model extracts the action response content, object feedback content, personnel feedback content, and action completion status from the response feedback data, and uses these contents for profile updates.

[0043] The target user service profile is updated based on user feedback, action response, and action completion status. User feedback reflects the target user's acceptance, rejection, inquiry, or supplementary needs regarding service actions; action response reflects operational feedback generated on the interactive terminal; and action completion status reflects whether the service action corresponding to the auxiliary action data has been advanced. The service personnel status profile is updated based on personnel feedback, action response, and action completion status. Personnel feedback reflects the service personnel's use, modification, skipping, or confirmation of auxiliary action data, used to adjust the service personnel's capability status, workload status, and response preference status.

[0044] The updated target object service profile is used to generate subsequent object service requirement data, the updated service personnel status profile is used to generate subsequent personnel service status data, and the action completion status is used to generate subsequent decision trigger data. The server-side control model generates subsequent service decision data based on the subsequent object service requirement data, subsequent personnel service status data, and subsequent decision trigger data. The subsequent service decision data is used to determine the subsequent target agent set, enabling the next round of auxiliary actions to reselect the agent's capability range based on feedback changes.

[0045] In scenarios where auxiliary action data includes multiple service reminders, a feedback status diversion approach is adopted to update the profile. In financial service scenarios, after the account manager's terminal receives customer follow-up reminders and product benefit explanations, the interactive terminal collects data on viewing, confirming, modifying scripts, and follow-up completion results. The server-side control model updates the target object's service profile based on the customer's response content and updates the service personnel's status profile based on the account manager's confirmation or skipping behavior, before generating subsequent service decision data.

[0046] When auxiliary action data includes interactive Q&A content, semantic feedback parsing is used to update the profile. In health management scenarios, after the health management terminal displays medication reminders or follow-up prompts to service personnel, the interactive terminal collects user responses, service personnel records, and follow-up completion results. The server-side control model extracts changes in health service needs from user responses, extracts changes in task load from service personnel records, and generates subsequent service decision data to determine the set of subsequent target agents.

[0047] In the fintech business, after the account manager sends a rights and benefits explanation, the financial customer responds by asking whether to continue the consultation or postpone the process. The server-side control model updates the financial customer service profile and the account manager status profile based on the feedback, and generates subsequent service decision data to determine the subsequent prompt-type or task-type intelligent agent.

[0048] In the field of healthcare, after a health management terminal sends a medication reminder, the user reports whether they have taken the medication, forgot to take it, or experienced discomfort. The server-side control model updates the user service profile and the health management personnel status profile based on the feedback, and generates subsequent service decision data to determine whether to use a follow-up or reminder-type intelligent agent.

[0049] This embodiment updates the target object service profile and service personnel status profile by responding to feedback data, and then generates subsequent service decision data based on the updated profile. The server-side control model can adjust the selection of subsequent intelligent agents according to the actual feedback after the assistance action, reduce the lag caused by fixed service actions, and improve the continuous updating capability of the service assistance process.

[0050] In one embodiment, step S10 includes: S101, establish a multimodal acquisition session on the interactive terminal, and configure a session identifier, modality source identifier, acquisition time identifier and terminal status identifier for the multimodal acquisition session; S102, receive multiple types of input segments based on the session identifier, and divide the multiple types of input segments into multiple modal segments among voice input segments, text input segments, image input segments and operation behavior segments according to the modal source identifier, and generate multimodal interactive input data; S103, by using the end-side multimodal model deployed on the interactive terminal, extracting multiple types of end-side perception segments corresponding to the multiple types of modal segments from the multimodal interactive input data, the multiple types of end-side perception segments include multiple types of segments among speech semantic segments, text semantic segments, image object segments and operation feature segments; S104, based on the acquisition time identifier and the modality source identifier, the multi-type end-side sensing segments are time-aligned and cross-modal associated through the end-side multimodal model, and acquisition status markers are added based on the terminal status identifier to generate cross-modal associated feature data; S105, through the terminal multimodal model, extract service intent candidate information, target object pointing information, personnel status prompt information and scene trigger prompt information from the cross-modal association feature data, and extract collection status information from the collection status marker; S106, the session identifier, modality source identifier, collection time identifier, service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information, and collection status information are encapsulated into structured perception feature data according to a field format that includes session field, modality field, time field, intent field, object field, personnel status field, scene trigger field, and collection status field.

[0051] In this embodiment, when an interactive terminal establishes a multimodal acquisition session, the terminal assigns a session identifier to each interaction and simultaneously configures a modality source identifier, an acquisition time identifier, and a terminal status identifier. The session identifier is used to merge different input content generated within the same interaction, preventing voice, text, images, and user actions from being sent to different processing flows. The modality source identifier indicates the source category of the input content, enabling the edge-side multimodal model to call the corresponding encoding branch. The acquisition time identifier records the generation order of the input content, allowing subsequent timing alignment to distinguish between sequential expressions under the same intent. The terminal status identifier records network status, terminal load, offline recognition status, and modality missing status, ensuring that the edge-side output retains differences in the acquisition environment.

[0052] After receiving multiple input segments based on session identifiers, the interactive terminal performs modality segmentation according to modality source identifiers. Voice input segments can be captured by a microphone; text input segments can be generated from input boxes, message windows, or form fields; image input segments can be generated by camera components, screenshot components, or file upload components; and operation behavior segments can be generated by terminal events such as page clicks, pauses, switches, confirmations, withdrawals, or closures. Multiple input segments do not need to simultaneously contain all modalities; as long as two or more types of input content exist under the same session identifier, they can be combined into multimodal interactive input data. During the combination process, the acquisition time identifier and modality source identifier are retained, enabling the on-device multimodal model to identify the temporal and source relationships between different segments.

[0053] The edge-side multimodal model is deployed on the interactive terminal side to perform preliminary semantic extraction. Speech input segments are processed through a speech encoding branch to generate speech semantic segments, text input segments through a text encoding branch to generate text semantic segments, image input segments through a visual encoding branch to generate image object segments, and operation behavior segments through a behavior encoding branch to generate operation feature segments. Speech semantic segments carry the intent expressed and object references within the speech content; text semantic segments carry service requests and supplementary information from the text input; image object segments carry objects, page regions, material categories, and recognizable text within the image; and operation feature segments carry the page state, objects of interest, and interaction preferences reflected in the terminal's operations.

[0054] The edge-side multimodal model performs temporal alignment and cross-modal association on various edge-side perception segments based on acquisition time identifiers and modality source identifiers. Temporal alignment is used to group speech semantic segments, text semantic segments, image object segments, and operation feature segments generated within the same time range into the same interaction context. Cross-modal association is used to merge content from different modalities that points to the same target object, the same page area, or the same service item. Terminal status identifiers are added to the acquisition status marker, so that the cross-modal association feature data simultaneously includes content recognition results and acquisition status results, avoiding subsequent processing from ignoring the impact of offline recognition, network fluctuations, or modality loss on the perception results.

[0055] The edge-side multimodal model extracts service intent candidate information, target object pointing information, personnel status prompts, and scene trigger prompts from cross-modal associated feature data, and extracts collection status information from collection status markers. Service intent candidate information indicates the possible service direction corresponding to the current input content; target object pointing information indicates the object involved in the current interaction; personnel status prompts indicate the operation or response status of the service personnel during the interaction; scene trigger prompts indicate the service scenario triggered by the current interaction; and collection status information indicates the status of the edge-side collection and recognition process. This information is written into the session field, modality field, time field, intent field, object field, personnel status field, scene trigger field, and collection status field, forming structured perceptual feature data.

[0056] This embodiment converts speech, text, images, and user actions into structured perceptual feature data with field meanings by performing session merging, modality segmentation, terminal-side semantic extraction, temporal alignment, and cross-modal association on multiple types of input fragments at the interactive terminal side. Since the structured perceptual feature data simultaneously contains service intent candidate information, target object pointing information, personnel status prompts, scene trigger prompts, and collection status information, subsequent processing can directly utilize the perception results generated at the terminal side. This reduces repeated parsing of scattered input content, minimizes recognition bias caused by fragmented multimodal data, and improves the completeness and usability of the perception results in the current interactive scenario.

[0057] In one embodiment, step S20 above includes: S201, the structured perception feature data is encapsulated into terminal-side transmission data, and the interactive terminal sends the terminal-side transmission data to the server-side control model; S202, the server-side control model parses the data sent from the terminal to obtain structured perception feature data, and extracts session identifier, target object pointing information, service intent candidate information, personnel status prompt information, scene trigger prompt information and collection status information from the structured perception feature data to generate service retrieval index data; S203, the server-side control model generates knowledge retrieval index data and object service retrieval index data based on the service retrieval index data; S204, through the server-side control model, based on the knowledge retrieval index data, retrieve scene knowledge data, organizational knowledge data, service resource data and operation guidance data from the knowledge resource storage area, and generate knowledge resource data based on the scene knowledge data, the organizational knowledge data, the service resource data and the operation guidance data; S205, through the server-side control model, based on the object service retrieval index data, the target object basic data, target object interaction record data, target object service status data, service personnel capability data, and service personnel task status data are retrieved from the object service data storage area, and the object service basic data is generated based on the target object basic data, the target object interaction record data, the target object service status data, the service personnel capability data, and the service personnel task status data; S206, through the server-side control model, the knowledge resource data and the object service basic data are bound to the structured perception feature data according to the session identifier.

[0058] In this embodiment, after the structured perception feature data is encapsulated on the terminal side, the interactive terminal organizes the structured perception feature data into terminal-side transmission data. The terminal-side transmission data may include a data body and transmission metadata. The data body carries session identifiers, target object information, service intent candidate information, personnel status prompts, scene trigger prompts, and collection status information. The transmission metadata carries the terminal source, transmission batch, field integrity status, and timestamp. When the interactive terminal sends the terminal-side transmission data to the server-side control model, it can use interface requests, message queue delivery, or terminal-cloud synchronization. Interface requests are suitable for immediate response, message queue delivery is suitable for concurrent uploads from multiple terminals, and terminal-cloud synchronization is suitable for temporary storage on the terminal followed by batch transmission.

[0059] After receiving data, the server-side control model parses the data body and transmission metadata to recover structured perception feature data. From this data, it extracts session identifiers, target object information, service intent candidate information, personnel status prompts, scene trigger prompts, and collection status information. Target object information limits the scope of objects to be retrieved; service intent candidate information limits the knowledge resource category; personnel status prompts limit the data scope on the service personnel side; scene trigger prompts limit the service scene; and collection status information marks the collection quality and modal completeness on the receiving end. The server-side control model combines this information into service retrieval index data, ensuring that resource retrieval is not dependent on a single keyword.

[0060] The server-side control model generates knowledge retrieval index data and object service retrieval index data based on service retrieval index data. The knowledge retrieval index data is oriented towards the knowledge resource storage area and includes index items such as scenario trigger prompts, service intent candidates, and service resource categories. It is used to retrieve scenario knowledge data, organizational knowledge data, service resource data, and operation guidance data. The object service retrieval index data is oriented towards the object service data storage area and includes index items such as target object pointers, personnel status prompts, and session identifiers. It is used to retrieve basic data of the target object, target object interaction record data, target object service status data, service personnel capability data, and service personnel task status data.

[0061] Knowledge resource data is retrieved from the knowledge resource storage area by the server-side control model based on the knowledge retrieval index data. Scenario knowledge data describes the classification and triggering conditions of service scenarios; organizational knowledge data provides referable business definitions and management requirements within the organization; service resource data provides callable resource content; and operation guidance data provides operation paths that service personnel can refer to. Object service basic data is retrieved from the object service data storage area by the server-side control model based on the object service retrieval index data. Target object basic data provides object attributes; target object interaction record data provides past interaction content; target object service status data provides service progress; service personnel capability data provides the scope of services that service personnel can handle; and service personnel task status data provides task occupancy status.

[0062] The server-side control model binds knowledge resource data, object service basic data, and structured perception feature data according to session identifiers, ensuring that all three types of data share the same session identifier. After binding, structured perception feature data provides the client-side perception results, knowledge resource data provides referenceable resources, and object service basic data provides the state source of objects and service personnel. All three types of data can be invoked within the same session scope, preventing data from different interaction rounds from being mixed in the same service judgment.

[0063] This embodiment encapsulates structured perception feature data into edge-side transmission data and sends it to the server-side control model. The server-side control model can extract target object pointers, service intent candidates, personnel status prompts, scene trigger prompts, and collection status from the field-based perception results, and generate retrieval indexes for different storage areas. After knowledge resource data and object service basic data are bound to structured perception feature data according to session identifiers, the server can obtain matching knowledge resources and object status data within the same session, reducing delays caused by manual cross-database retrieval and data mixing, improving resource retrieval accuracy and the continuity of subsequent service processing.

[0064] In one embodiment, step S30 above includes: S301, through the server-side control model, read the mutually bound structured perception feature data, knowledge resource data and object service basic data according to the session identifier; S302, through the server-side control model, extract service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information and collection status information from the structured perception feature data; S303, through the server-side control model, scene semantic data is generated based on the scene trigger prompt information and the knowledge resource data, and resource boundary data is generated based on the knowledge resource data; S304, Service intent data is generated based on the service intent candidate information, the scene semantic data, the resource boundary data, and the collection status information through the server-side control model; S305, through the server-side control model, target object status data is extracted from the object service basic data based on the target object pointing information, and a target object service profile is generated based on the target object status data, the service intent data and the knowledge resource data; S306, through the server-side control model, service personnel status data is extracted from the object service basic data based on the personnel status prompt information, and a service personnel status profile is generated based on the service personnel status data, the service intent data, and the knowledge resource data.

[0065] In this embodiment, the server-side control model reads the already bound structured perception feature data, knowledge resource data, and object service basic data according to the session identifier. The session identifier is used to limit the scope of the same interaction and avoid the mixing of data generated from different interaction rounds. The structured perception feature data provides the perception results after edge recognition, the knowledge resource data provides scene knowledge, service resources, and operation guidance, and the object service basic data provides the basic status of the target object and service personnel. During the reading process, the server-side control model can establish a data mapping table according to the session identifier, map the three types of data to the same session node, and mark the status of missing fields.

[0066] The server-side control model extracts service intent candidate information, target object pointing information, personnel status prompts, scene trigger prompts, and collection status information from structured perception feature data. Service intent candidate information represents the service direction identified by the client-side; target object pointing information locates the object involved in the current interaction; personnel status prompts reflect the operational or response status of service personnel; scene trigger prompts identify the current service scene; and collection status information reflects the data source, modal completeness, and recognition status of the client-side data. The server-side control model can convert the above information into structured input usable in subsequent generation processes through field parsing, field normalization, and field confidence labeling.

[0067] The server-side control model generates scene semantic data based on scene trigger prompts and knowledge resource data. During generation, the scene trigger prompts are matched with scene knowledge in the knowledge resource data to extract the scene category, scene conditions, and scene processing scope corresponding to the current service scene. The server-side control model also generates resource boundary data based on the knowledge resource data. This resource boundary data can include the scope of callable resources, operational restrictions, service resource types, and the scope of prompt content. This data is used to constrain the formation of subsequent service intent data and reduce inconsistencies between service intents and available resources.

[0068] The server-side control model generates service intent data based on service intent candidate information, scene semantic data, resource boundary data, and acquisition status information. Service intent candidate information provides preliminary identification results on the client side; scene semantic data is used to refine the service scenario; resource boundary data is used to limit the possible service directions; and acquisition status information is used to indicate the usability of the client-side identification results. Service intent data can include service category, service goal, triggering scenario, urgency level, and manageable scope, ensuring that service intents are not derived from a single input but are formed in combination with scene and resource conditions.

[0069] The server-side control model extracts target object status data from the object service basic data based on the target object's pointing information. Target object status data can include basic object attributes, interaction status, service progress, demand status, and unfinished service items. The server-side control model then combines the target object status data, service intent data, and knowledge resource data to generate a target object service profile. This profile can include object activity status, demand status, service gap status, and resource matching status, reflecting the target object's service status within the current session scope.

[0070] The server-side control model extracts service personnel status data from the object service's basic data based on personnel status prompts. This data can include capability status, task load, service proficiency status, response preferences, and current task occupancy. The server-side control model then combines this service personnel status data, service intent data, and knowledge resource data to generate a service personnel status profile. This profile can include personnel capability status, task load status, response preference status, and service readiness status, reflecting the service personnel's ability to accept the current service intent.

[0071] This embodiment reads and associates structured perception feature data, knowledge resource data, and object service basic data through session identifiers. The server-side control model can integrate edge-side perception results, callable knowledge content, and the basic status of objects and service personnel within the same session. Service intent candidate information is corrected by scene semantic data, resource boundary data, and collection status information to generate service intent data. Target object pointing information and personnel status prompt information drive the generation of target object service profiles and service personnel status profiles, respectively, thereby reducing judgment bias caused by single perception results and improving the consistency between service intent data, target object service profiles, and service personnel status profiles.

[0072] In one embodiment, step S40 above includes: S401, through the server-side control model, extract service target information, trigger scenario information and service urgency information from the service intent data, extract object activity status data, object demand status data and object service gap data from the target object service profile, and extract personnel capability status data, personnel task load data and personnel response preference data from the service personnel status profile; S402, through the server-side control model, service constraint data is generated based on the trigger scenario information, the object demand status data, the object service gap data, the personnel capability status data, and the personnel task load data; S403, Based on the service target information and the service constraint data, the server-side control model generates service action candidate data; S404, Through the server-side control model, service priority data is generated based on the service urgency information, the object activity status data, and the personnel task load data; S405, through the server-side control model, generate agent invocation conditions and auxiliary action output conditions based on the service action candidate data, the service priority data, and the personnel response preference data, and encapsulate the service action candidate data, the service priority data, the agent invocation conditions, and the auxiliary action output conditions into service decision data.

[0073] In this embodiment, the server-side control model extracts service target information, trigger scenario information, and service urgency information from the service intent data. Service target information identifies the direction of the service action to be generated, trigger scenario information identifies the scenario type in which the service action occurs, and service urgency information identifies the processing priority of the service action in terms of time. During the extraction process, the server-side control model can parse the category field, target field, scenario field, and time sequence field in the service intent data and convert the different fields into a unified decision input format.

[0074] The server-side control model extracts object activity status data, object demand status data, and object service gap data from the target object service profile. Object activity status data reflects the target object's recent service response, interaction frequency, and status changes; object demand status data reflects the service directions that the target object currently needs to have met; and object service gap data reflects service items that the target object has not yet completed, reached, or responded to. The server-side control model extracts personnel capability status data, personnel workload data, and personnel response preference data from the service personnel status profile. Personnel capability status data reflects the range of service actions that service personnel can undertake; personnel workload data reflects the current service task occupancy; and personnel response preference data reflects the service personnel's response tendencies to prompt methods, task types, and interaction methods.

[0075] The server-side control model generates service constraint data based on trigger scenario information, object requirement status data, object service gap data, personnel capability status data, and personnel task load data. This service constraint data limits the scope of subsequent service actions, excluding actions that are incompatible with the current scenario, inconsistent with object requirements, exceed personnel capabilities, or are unsuitable for the current task load. Service constraint data can include scenario constraints, requirement constraints, capability constraints, and load constraints, ensuring that candidate service action data is subject to multi-dimensional state limitations before generation.

[0076] The server-side control model generates candidate service actions based on service target information and service constraint data. Service target information provides the direction for action generation, while service constraint data limits the scope of action availability. The combination of these two elements forms candidate actions that can be invoked by subsequent agents. Candidate service actions can include information prompts, task generation actions, resource invocation actions, interactive assistance actions, and reminder triggering actions. The server-side control model then generates service priority data based on service urgency information, object activity status data, and personnel workload data. Service urgency information reflects the time sensitivity of service actions, object activity status data reflects the likelihood of target object response, and personnel workload data reflects the capacity of service personnel to handle the workload. These three types of data jointly determine the order of different candidate actions.

[0077] The server-side control model generates agent invocation conditions and auxiliary action output conditions based on service action candidate data, service priority data, and user response preference data. Agent invocation conditions indicate the type of agent capability, action processing type, and output type to be invoked subsequently; auxiliary action output conditions indicate the display format, triggering method, and interaction entry point of the auxiliary action on the interactive terminal. The server-side control model encapsulates the service action candidate data, service priority data, agent invocation conditions, and auxiliary action output conditions into service decision data, enabling the service decision data to simultaneously carry the action content, action order, agent invocation basis, and terminal output basis.

[0078] This embodiment extracts the necessary state data for decision-making from service intent data, target object service profiles, and service personnel status profiles through a server-side control model. It also transforms scenario, object needs, object gaps, personnel capabilities, and task load into service constraint data, reducing the bias caused by directly generating actions from a single intent. The service action candidate data, combined with service priority data and personnel response preference data, is encapsulated into service decision data containing agent invocation conditions and auxiliary action output conditions. This provides a clear basis for subsequent agent selection and auxiliary action output, improving the matching between service decisions and the current service state.

[0079] In one embodiment, step S50 above includes: S501, through the server-side control model, extract the agent invocation conditions, service action candidate data, service priority data and auxiliary action output conditions from the service decision data; S502, through the server-side control model, candidate intelligent agents are selected from the intelligent agent registration data based on the intelligent agent invocation conditions, and the action capabilities of the candidate intelligent agents are compared with the service action candidate data to generate candidate intelligent agent capability data; S503, through the server-side control model, based on the service priority data and the candidate agent capability data, determine the target agent set from the candidate agents; S504, through the server-side control model, the service action candidate data and the auxiliary action output conditions are distributed to the target intelligent agent set, and an action fragment set corresponding to the target intelligent agent set is generated through the target intelligent agent set. The action fragment set includes any one or more types of fragments among task action fragments, resource action fragments and interactive prompt fragments. S505, through the server-side control model, based on the service priority data, conflict resolution and presentation order arrangement are performed on the action fragment set to generate auxiliary action data; S506, through the server-side control model, the auxiliary action data is converted into terminal display data and terminal interaction trigger data based on the auxiliary action output conditions, and the terminal display data and terminal interaction trigger data are sent to the interactive terminal.

[0080] In this embodiment, the server-side control model extracts agent invocation conditions, service action candidate data, service priority data, and auxiliary action output conditions from the service decision data. Agent invocation conditions limit the scope of capabilities of the agent to be invoked, such as task generation capabilities, resource extraction capabilities, interactive prompt capabilities, or page triggering capabilities. Service action candidate data carries the content of generateable service actions, such as reminders, recommendations, Q&A, tasks, prompts, or resource displays. Service priority data marks the processing order of different service actions. Auxiliary action output conditions limit the presentation and triggering methods of auxiliary action data in the interactive terminal, such as card display, pop-up prompts, session insertion, task entry generation, or page jump triggering.

[0081] Agent registration data records the capability tags, input data types, output data types, action types that can be processed, and running status of callable agents. The server-side control model filters candidate agents from the agent registration data based on agent invocation conditions, and then compares the capability tags and action types of the candidate agents with the service action candidate data. Action capability comparison can be performed based on action type consistency, input field matching degree, output data type, and running status to generate candidate agent capability data. The candidate agent capability data indicates whether the candidate agent can undertake the current service action, and the output type when undertaking different service actions.

[0082] The server-side control model determines the target agent set based on service priority data and candidate agent capability data. Service priority data is used to select actions that require more advance generation, while candidate agent capability data is used to select agents capable of generating the corresponding actions. The target agent set can contain a single agent or multiple agents. When multiple agents are selected, the server-side control model distributes service action candidate data and auxiliary action output conditions to different agents, enabling the generation of task-related, resource-related, and prompt-related content respectively.

[0083] The target intelligent agent set generates a set of action fragments. Task action fragments can include to-do items, follow-up reminders, or service task entry points. Resource action fragments can include resource summaries, resource entry points, knowledge fragments, or material prompts. Interaction prompt fragments can include communication scripts, Q&A guidance, reminder text, or page interaction suggestions. The server-side control model resolves conflicts and arranges the presentation order of the action fragment set based on service priority data. Conflict resolution can handle situations such as duplicate reminders, repeated display of the same resource, contradictory task orders, or inconsistent prompt content. The presentation order arrangement determines the display order and triggering sequence of action fragments in the interactive terminal.

[0084] The server-side control model converts auxiliary action data into terminal display data and terminal interaction trigger data based on the auxiliary action output conditions. Terminal display data is used to create visible content on the interactive terminal, such as reminder cards, service task lists, resource summaries, or communication prompts. Terminal interaction trigger data is used to create actionable entry points on the interactive terminal, such as jump entry points, confirmation buttons, Q&A entry points, or task processing entry points. After the terminal display data and terminal interaction trigger data are sent to the interactive terminal, the auxiliary action data can be presented in a form that service personnel can view and operate.

[0085] This embodiment extracts agent invocation conditions, service action candidate data, service priority data, and auxiliary action output conditions from service decision data through a server-side control model. Based on agent registration data, it filters the target agent set, enabling service actions to be generated by agents with corresponding capabilities. After the target agent set generates a set of action fragments, the server-side control model resolves conflicts and arranges the presentation order according to the service priority data. Then, based on the auxiliary action output conditions, it converts the data into terminal display data and terminal interaction trigger data. This ensures that the auxiliary action data is sent to the interactive terminal according to the service action content, processing order, and terminal presentation requirements, reducing the time spent manually selecting tools and organizing prompts, and improving the consistency between auxiliary action generation and terminal presentation.

[0086] In one embodiment, step S60 above includes: S601, collect terminal feedback content, object feedback content and personnel feedback content formed in response to the auxiliary action data through the interactive terminal, and generate response feedback data based on the terminal feedback content, the object feedback content and the personnel feedback content; S602, through the server-side control model, extract the action response content, object feedback content, personnel feedback content and action completion status from the response feedback data; S603, through the server-side control model, target object profile update data is generated based on the object feedback content, the action response content and the action completion status, and the target object service profile is updated based on the target object profile update data to generate the updated target object service profile; S604, through the server-side control model, service personnel profile update data is generated based on the personnel feedback content, the action response content, and the action completion status, and the service personnel status profile is updated based on the service personnel profile update data to generate the updated service personnel status profile; S605, through the server-side control model, generate subsequent object service requirement data based on the updated target object service profile, generate subsequent personnel service status data based on the updated service personnel status profile, and generate subsequent decision trigger data based on the action completion status. S606, Through the server-side control model, subsequent service decision data for determining the subsequent target intelligent agent set is generated based on the subsequent object service demand data, the subsequent personnel service status data, and the subsequent decision trigger data.

[0087] In this embodiment, after the auxiliary action data arrives, the interactive terminal collects terminal feedback content, object feedback content, and personnel feedback content corresponding to the auxiliary action data. Terminal feedback content may include the display status, click status, dwell time, confirmation status, closed status, and redirect status of the auxiliary action data. Object feedback content may include the object's response, confirmation, rejection, supplementary explanation, and further inquiry regarding the reminder, recommendation, Q&A, or task content. Personnel feedback content may include service personnel's viewing, acceptance, modification, skipping, reassignment, and completion confirmation of the auxiliary action data. The interactive terminal aggregates the three types of feedback content according to the same auxiliary action data identifier and session identifier, generating response feedback data so that subsequent profile updates can distinguish the feedback source and the feedback object.

[0088] The server-side control model extracts action response content, object feedback content, personnel feedback content, and action completion status from the response feedback data. Action response content records the interaction results generated by the interactive terminal around the auxiliary action data, including actions such as clicking, confirming, editing, assigning, and closing. Object feedback content records the target object's reaction to the service action corresponding to the auxiliary action data. Personnel feedback content records how service personnel use the auxiliary action data. Action completion status records whether the task, reminder, communication, or resource push corresponding to the auxiliary action data has been completed, and whether the completion result meets the triggering conditions for subsequent decisions.

[0089] The target object profile update data is generated jointly by the object's feedback content, action response content, and action completion status. Object feedback content reflects the target object's changing needs and acceptance level regarding service actions; action response content reflects the interaction results related to the object's service on the interactive terminal; and action completion status reflects the progress of the service action. The server-side control model writes the target object profile update data into fields such as demand status, service gaps, activity status, and resource preferences in the target object service profile, generating the updated target object service profile.

[0090] Service personnel profile update data is generated jointly from personnel feedback, action response content, and action completion status. Personnel feedback reflects the service personnel's adoption and modification of auxiliary action data; action response content reflects the service personnel's operational results on the interactive terminal; and action completion status reflects the service personnel's progress in advancing the service action. The server-side control model writes the updated service personnel profile data into fields such as capability status, task load, response preferences, and service readiness status in the service personnel status profile, generating the updated service personnel status profile.

[0091] The updated target object service profile is used to generate subsequent object service requirement data. This data reflects the target object's new needs, unfinished tasks, changes in service preferences, and continued outreach needs after feedback. The updated service personnel status profile is used to generate subsequent personnel service status data. This data reflects changes in service personnel's workload, capacity, prompting preferences, and service readiness status after feedback. The server-side control model also generates subsequent decision trigger data based on action completion status. This data identifies whether further prompting is needed, whether to switch to another agent, postpone action generation, or generate a new service action.

[0092] The server-side control model generates subsequent service decision data based on subsequent object service demand data, subsequent personnel service status data, and subsequent decision trigger data. Subsequent object service demand data provides the results of changes on the object side, subsequent personnel service status data provides the results of changes on the personnel side, and subsequent decision trigger data provides the basis for deciding whether to continue generating actions. The subsequent service decision data is used to determine the set of subsequent target agents, enabling the selection of subsequent agents to be adjusted based on the feedback profile status and service action completion status.

[0093] This embodiment generates response feedback data by collecting feedback from terminals, objects, and personnel. The server-side control model then updates the target object service profile and the service personnel status profile, respectively. The actual feedback after the auxiliary action data is generated can be transformed into the basis for profile updates. The updated target object service profile and the updated service personnel status profile further generate subsequent object service demand data and subsequent personnel service status data. Combined with the action completion status, subsequent service decision data is generated. This allows the determination of the subsequent target intelligent agent set to be adjusted based on the feedback of object needs, personnel status, and action completion results, reducing service lag caused by fixed action patterns and improving the continuous update capability of the service assistance process.

[0094] In one embodiment, in a fintech business scenario, an insurance service institution deploys an object-oriented service auxiliary application for its agents, with the interactive terminal being a mobile sales terminal used by the agents. When an agent visits a financial customer, the interactive terminal establishes a multimodal data collection session and configures a session identifier, modality source identifier, collection time identifier, and terminal status identifier for the visit. When the agent communicates with the financial customer about policy benefits, product upgrades, family protection gaps, and claims services, the interactive terminal collects voice input segments; when the agent enters customer concerns in the customer's notes section, the interactive terminal collects text input segments; when the agent photographs the policy page, protection plan page, and product benefits page authorized by the customer, the interactive terminal collects image input segments; when the agent views today's business updates, switches product description pages, clicks the benefits description entry, and closes the risk warning card in the application, the interactive terminal collects action behavior segments. The interactive terminal categorizes these segments into the same visit process according to the session identifier and classifies them into multiple modal segments, including voice input segments, text input segments, image input segments, and action behavior segments, according to the modality source identifier, generating multimodal interactive input data. After reading the multimodal interactive input data, the edge-side multimodal model deployed on the interactive terminal extracts the customer's voice semantic segments regarding protection gaps and rights claims from the voice input segments through the voice encoding branch; extracts the customer's text semantic segments regarding family protection and premium budget from the notes through the text encoding branch; extracts the product name, rights entry, scope of liability, and page text to form image object segments from the policy page and rights page through the image encoding branch; and extracts operation feature segments from page switching, dwell, and click operations through the behavior encoding branch. The edge-side multimodal model then performs temporal alignment and cross-modal association of the voice semantic segments, text semantic segments, image object segments, and operation feature segments based on the collection time identifier and modality source identifier, and adds a collection status marker according to the terminal status identifier to generate cross-modal associated feature data. The edge-side multimodal model extracts service intent candidate information, target object pointing information, personnel status prompt information, and scene trigger prompt information from cross-modal associated feature data, and extracts collection status information from collection status markers. Then, it writes the session identifier, modality source identifier, collection time identifier, service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information, and collection status information into the session field, modality field, time field, intent field, object field, personnel status field, scene trigger field, and collection status field, and encapsulates them into structured perception feature data.

[0095] After the structured perception feature data is formed, the interactive terminal encapsulates it into client-side transmission data and sends it to the server-side control model. The server-side control model parses the client-side transmission data to obtain the structured perception feature data, and extracts session identifiers, target object information, service intent candidate information, personnel status prompts, scenario trigger prompts, and collection status information from the structured perception feature data to generate service retrieval index data. Based on the service retrieval index data, the server-side control model generates knowledge retrieval index data and object service retrieval index data. The knowledge retrieval index data is used to access the knowledge resource storage area to retrieve scenario knowledge data, organizational knowledge data, service resource data, and operation guidance data. Scenario knowledge data may include business scenario knowledge such as face-to-face visits, rights and benefits claims, policy review, customer reactivation, and product descriptions; organizational knowledge data may include company policies, compliance notices, risk communication guidelines, and service specifications; service resource data may include product descriptions, rights and benefits descriptions, business updates, industry hot topics, and event information; and operation guidance data may include application operation manuals, rights and benefits claims guidelines, and face-to-face visit material retrieval guidelines. The server-side control model generates knowledge resource data based on scenario knowledge data, organizational knowledge data, service resource data, and operation guidance data. Object service retrieval index data is used to access the object service data storage area, retrieving target object basic data, target object interaction record data, target object service status data, service personnel capability data, and service personnel task status data. Target object basic data may include the financial customer's policy overview, membership level, rights holding status, and authorized service attributes; target object interaction record data may include upcoming pending matters, past face-to-face visit records, telephone communication records, and application outreach records; target object service status data may include rights claim status, product description status, risk warning status, and dormant customer reactivation status; service personnel capability data may include the specialist's product types of expertise, communication ability tags, and AI coaching training results; service personnel task status data may include today's face-to-face visit arrangements, tasks to be followed up, and current task load. The server-side control model generates object service basic data based on the above data and binds knowledge resource data, object service basic data, and structured perception feature data according to session identifiers, ensuring that all three types of data have the same session identifier.

[0096] The server-side control model reads the interlinked structured perception feature data, knowledge resource data, and object service basic data according to the session identifier. It extracts service intent candidate information, target object information, personnel status prompts, scenario trigger prompts, and collection status information from the structured perception feature data. Based on the scenario trigger prompts and knowledge resource data, the server-side control model generates scenario semantic data. For example, it identifies the current face-to-face visit as a scenario for customer rights review and gap explanation, and generates resource boundary data based on the knowledge resource data, such as limiting the available product descriptions, rights explanations, risk warnings, AI coaching scripts, and business updates. The server-side control model generates service intent data based on the service intent candidate information, scenario semantic data, resource boundary data, and collection status information. This service intent data can indicate that assistance is currently needed to help the specialist complete customer rights explanations, gap explanations, and subsequent follow-up arrangements. Based on the target object information, the server-side control model extracts target object status data from the object service basic data and combines it with the service intent data and knowledge resource data to generate a target object service profile. The target customer service profile can include customer activity status, customer demand status, service gap status, and resource matching status. For example, a customer who has recently viewed the critical illness insurance page multiple times but has not completed the claim process is a target customer who can be provided with benefits explanations and insurance supplementation reminders. The server-side control model extracts service personnel status data from the target service basic data based on personnel status prompts and combines this with service intent data and knowledge resource data to generate a service personnel status profile. The service personnel status profile can include specialist capability status, workload status, response preference status, and service readiness status. For example, a specialist may currently have many face-to-face visits but has good training records in explaining insurance products and customer communication techniques.

[0097] The server-side control model generates service decision data based on service intent data, target object service profiles, and service personnel status profiles. From the service intent data, the model extracts service target information, trigger scenario information, and service urgency information; from the target object service profile, it extracts object activity status data, object demand status data, and object service gap data; and from the service personnel status profile, it extracts personnel capability status data, personnel workload data, and personnel response preference data. Based on trigger scenario information, target demand status data, target service gap data, personnel capability status data, and personnel workload data, the model generates service constraint data. For example, during a face-to-face visit, priority is given to explaining rights and benefits and addressing gaps in protection, avoiding the promotion of product content unrelated to the customer's current needs. Based on service target information and service constraint data, the model generates candidate service actions, such as explaining rights and benefits, risk warnings, customer follow-up, AI-assisted training scripts, and service resource recommendations. Based on service urgency information, target activity status data, and personnel workload data, the model generates service priority data, such as prioritizing explanations of rights and benefits and risk warnings, and scheduling follow-up tasks to be triggered after the face-to-face visit. The server-side control model then generates agent invocation conditions and auxiliary action output conditions based on service action candidate data, service priority data, and personnel response preference data, and encapsulates the service action candidate data, service priority data, agent invocation conditions, and auxiliary action output conditions into service decision data.

[0098] The server-side control model extracts agent invocation conditions, service action candidate data, service priority data, and auxiliary action output conditions from service decision data. Based on the agent invocation conditions, it filters candidate agents from the agent registration data and compares their action capabilities with the service action candidate data to generate candidate agent capability data. Based on the service priority data and candidate agent capability data, the server-side control model determines the target agent set from the candidate agents. The target agent set can include task generation agents, resource extraction agents, and interactive prompt agents. The server-side control model distributes the service action candidate data and auxiliary action output conditions to the target agent set. Task generation agents generate customer follow-up tasks, post-face-visit to-do items, and dormant customer reactivation tasks; resource extraction agents extract resource summaries from today's business updates, company policies, product descriptions, and benefit descriptions; interactive prompt agents generate communication scripts, risk warnings, AI-assisted Q&A, and suggestions for responding to customer objections. After the target intelligent agent set generates a set of action fragments, the server-side control model resolves conflicts and arranges the presentation order of the action fragments based on service priority data. This avoids the repetition of the same benefit statement and conflicts between the display order of risk warnings and product recommendations, and generates auxiliary action data. Based on the auxiliary action output conditions, the server-side control model converts the auxiliary action data into terminal display data and terminal interaction trigger data, and then sends them to the interactive terminal. The interactive terminal can display "Today's Business Express Card," "Customer Benefit Statement Card," "AI-Assisted Communication Script Card," "Follow-up Task Entry," "Customer Recall Reminder," and "Device Status Reminder" to specialists, enabling them to obtain timely auxiliary information during face-to-face visits.

[0099] After auxiliary action data is sent to the interactive terminal, the terminal collects terminal feedback, target feedback, and personnel feedback based on the auxiliary action data, and generates response feedback data based on these three types of feedback. Terminal feedback may include instances such as a specialist viewing a benefits description card, clicking on the AI-assisted Q&A entry, confirming a follow-up task, turning off device status reminders, and remaining on the risk warning page. Target feedback may include instances such as a financial customer accepting a benefits claim suggestion, requesting supplementary consultation, rejecting the current product description, or requesting follow-up contact. Personnel feedback may include instances such as a specialist adopting or modifying a sales script, skipping a task, completing a face-to-face interview, or transitioning to team collaboration. The server-side control model extracts action response content, target feedback content, personnel feedback content, and action completion status from the response feedback data. Based on the target feedback content, action response content, and action completion status, the server-side control model generates target target profile update data and updates the target target service profile, generating an updated target target service profile. For example, when a customer accepts a benefits description and schedules follow-up communication, the customer's activity status and service progress are updated; when a customer rejects the current recommendation but focuses on family protection, the demand status and service gap status are adjusted. The server-side control model generates updated service personnel profile data based on personnel feedback, action response content, and action completion status. It then updates the service personnel status profile based on this updated data, generating a new service personnel status profile. For example, when a specialist adopts the AI-assisted practice script and completes a follow-up task, the service preparation status and task completion status are updated; when a specialist repeatedly skips similar prompts, the personnel response preference status is updated. The server-side control model then generates subsequent object service requirement data based on the updated target object service profile, subsequent personnel service status data based on the updated service personnel status profile, and subsequent decision trigger data based on the action completion status. Based on the subsequent object service requirement data, subsequent personnel service status data, and subsequent decision trigger data, the server-side control model generates subsequent service decision data to determine the set of subsequent target agents. This allows for the selection of task-generating agents, resource-extraction agents, or interactive prompt agents in the next round of service, forming a continuously updated object service assistance process.

[0100] In one embodiment, a multimodal interaction-based object service assistance device is provided, which corresponds one-to-one with the multimodal interaction-based object service assistance method described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the object service assistance device based on multimodal interaction of the present invention. The object service assistance device based on multimodal interaction includes: a terminal-side multimodal perception module 10, a server-side resource retrieval module 20, an intent profile generation module 30, a service decision generation module 40, an agent action orchestration module 50, and a feedback update module 60. Detailed descriptions of each functional module are as follows: The edge-side multimodal perception module 10 is used to acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the edge-side multimodal model; The server-side resource retrieval module 20 is used to send the structured perception feature data to the server-side control model, and retrieve knowledge resource data and object service basic data through the server-side control model; The intent profile generation module 30 is used to generate service intent data, target object service profile, and service personnel status profile based on the structured perception feature data, the knowledge resource data, and the object service basic data through the server-side control model. The service decision generation module 40 is used to generate service decision data based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. The intelligent agent action orchestration module 50 is used to determine a target intelligent agent set based on the service decision data, generate auxiliary action data through the target intelligent agent set, and send the auxiliary action data to the interactive terminal; The feedback update module 60 is used to collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

[0101] In one embodiment, the end-side multimodal sensing module 10 is specifically used for: Establish a multimodal acquisition session on the interactive terminal, and configure a session identifier, modality source identifier, acquisition time identifier, and terminal status identifier for the multimodal acquisition session; Based on the session identifier, multiple types of input segments are received, and according to the modality source identifier, the multiple types of input segments are divided into multiple modal segments among voice input segments, text input segments, image input segments and operation behavior segments, to generate multimodal interactive input data; By deploying a multimodal model on the interactive terminal, multiple types of end-side perception segments corresponding to the multiple modal segments are extracted from the multimodal interactive input data. The multiple types of end-side perception segments include multiple segments from speech semantic segments, text semantic segments, image object segments, and operation feature segments. Based on the acquisition time identifier and the modality source identifier, the multi-type end-side sensing segments are time-aligned and cross-modal associated using the end-side multimodal model, and acquisition status markers are added based on the terminal status identifier to generate cross-modal associated feature data; Through the edge multimodal model, service intent candidate information, target object pointing information, personnel status prompt information and scene trigger prompt information are extracted from the cross-modal association feature data, and collection status information is extracted from the collection status marker. The session identifier, modality source identifier, collection time identifier, service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information, and collection status information are encapsulated into structured perception feature data according to a field format that includes session field, modality field, time field, intent field, object field, personnel status field, scene trigger field, and collection status field.

[0102] In one embodiment, the server-side resource retrieval module 20 is specifically used for: The structured perception feature data is encapsulated into edge-side transmission data, and the edge-side transmission data is sent to the server-side control model by the interactive terminal. The server-side control model parses the data sent from the terminal to obtain structured perception feature data, and extracts session identifier, target object pointing information, service intent candidate information, personnel status prompt information, scene trigger prompt information and collection status information from the structured perception feature data to generate service retrieval index data. The server-side control model generates knowledge retrieval index data and object service retrieval index data based on the service retrieval index data. Through the server-side control model, scenario knowledge data, organizational knowledge data, service resource data, and operation guidance data are retrieved from the knowledge resource storage area based on the knowledge retrieval index data, and knowledge resource data is generated based on the scenario knowledge data, the organizational knowledge data, the service resource data, and the operation guidance data; Through the server-side control model, the target object basic data, target object interaction record data, target object service status data, service personnel capability data, and service personnel task status data are retrieved from the object service data storage area based on the object service retrieval index data. Based on the target object basic data, the target object interaction record data, the target object service status data, the service personnel capability data, and the service personnel task status data, the target service basic data is generated. The knowledge resource data and the object service basic data are bound to the structured perception feature data according to the session identifier through the server control model.

[0103] In one embodiment, the intent profile generation module 30 is specifically used for: Through the server-side control model, the mutually bound structured perception feature data, knowledge resource data, and object service basic data are read according to the session identifier; Through the server-side control model, service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information, and collection status information are extracted from the structured perception feature data; Through the server-side control model, scene semantic data is generated based on the scene trigger prompt information and the knowledge resource data, and resource boundary data is generated based on the knowledge resource data; Service intent data is generated based on the service intent candidate information, the scene semantic data, the resource boundary data, and the collection status information through the server-side control model. Through the server-side control model, target object status data is extracted from the object service basic data based on the target object pointing information, and a target object service profile is generated based on the target object status data, the service intent data, and the knowledge resource data. Through the server-side control model, service personnel status data is extracted from the object service basic data based on the personnel status prompt information, and a service personnel status profile is generated based on the service personnel status data, the service intent data, and the knowledge resource data.

[0104] In one embodiment, the service decision generation module 40 is specifically used for: Through the server-side control model, service target information, trigger scenario information and service urgency information are extracted from the service intent data; object activity status data, object demand status data and object service gap data are extracted from the target object service profile; and personnel capability status data, personnel task load data and personnel response preference data are extracted from the service personnel status profile. Service constraint data is generated based on the trigger scenario information, object demand status data, object service gap data, personnel capability status data, and personnel task load data through the server-side control model. Based on the service target information and the service constraint data, the server-side control model generates candidate data for service actions. The server-side control model generates service priority data based on the service urgency information, the object activity status data, and the personnel task load data. The server-side control model generates agent invocation conditions and auxiliary action output conditions based on the service action candidate data, service priority data, and personnel response preference data, and encapsulates the service action candidate data, service priority data, agent invocation conditions, and auxiliary action output conditions into service decision data.

[0105] In one embodiment, the agent action orchestration module 50 is specifically used for: Through the server-side control model, intelligent agent invocation conditions, service action candidate data, service priority data, and auxiliary action output conditions are extracted from the service decision data. Through the server-side control model, candidate agents are selected from the agent registration data based on the agent invocation conditions, and the action capabilities of the candidate agents are compared with the service action candidate data to generate candidate agent capability data. Based on the service priority data and the candidate agent capability data, the target agent set is determined from the candidate agents using the server-side control model. The server-side control model distributes the service action candidate data and the auxiliary action output conditions to the target intelligent agent set, and generates an action fragment set corresponding to the target intelligent agent set. The action fragment set includes any one or more types of fragments among task action fragments, resource action fragments, and interactive prompt fragments. Through the server-side control model, conflict resolution and presentation order arrangement are performed on the action fragment set based on the service priority data to generate auxiliary action data. The server-side control model converts the auxiliary action data into terminal display data and terminal interaction trigger data based on the auxiliary action output conditions, and then sends the terminal display data and terminal interaction trigger data to the interactive terminal.

[0106] In one embodiment, the feedback update module 60 is specifically used for: The system collects terminal feedback content, object feedback content, and personnel feedback content in response to the auxiliary action data through an interactive terminal, and generates response feedback data based on the terminal feedback content, object feedback content, and personnel feedback content. The server-side control model extracts action response content, object feedback content, personnel feedback content, and action completion status from the response feedback data. Through the server-side control model, target object profile update data is generated based on the object feedback content, the action response content, and the action completion status. The target object service profile is then updated based on the target object profile update data to generate the updated target object service profile. The server-side control model generates updated service personnel profile data based on the personnel feedback content, the action response content, and the action completion status, and updates the service personnel status profile based on the updated service personnel profile data, thus generating an updated service personnel status profile. Through the server-side control model, subsequent object service requirement data is generated based on the updated target object service profile, subsequent personnel service status data is generated based on the updated service personnel status profile, and subsequent decision triggering data is generated based on the action completion status. The server-side control model generates subsequent service decision data for determining the set of subsequent target intelligent agents based on the subsequent object service demand data, the subsequent personnel service status data, and the subsequent decision triggering data.

[0107] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multimodal interaction-based object service auxiliary method on the server side.

[0108] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a multimodal interaction-based object service assistance method.

[0109] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the terminal-side multimodal model; The structured perception feature data is sent to the server-side control model, and the knowledge resource data and object service basic data are retrieved through the server-side control model. Through the server-side control model, service intent data, target object service profiles, and service personnel status profiles are generated based on the structured perception feature data, the knowledge resource data, and the object service basic data. Service decision data is generated based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. Based on the service decision data, a set of target intelligent agents is determined, auxiliary action data is generated through the set of target intelligent agents, and the auxiliary action data is sent to the interactive terminal; Collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

[0110] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the terminal-side multimodal model; The structured perception feature data is sent to the server-side control model, and the knowledge resource data and object service basic data are retrieved through the server-side control model. Through the server-side control model, service intent data, target object service profiles, and service personnel status profiles are generated based on the structured perception feature data, the knowledge resource data, and the object service basic data. Service decision data is generated based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. Based on the service decision data, a set of target intelligent agents is determined, auxiliary action data is generated through the set of target intelligent agents, and the auxiliary action data is sent to the interactive terminal; Collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

[0111] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0112] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0113] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0114] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0115] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for object service assistance based on multi-modal interaction, the method comprising: Includes the following steps: Acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the terminal-side multimodal model; The structured perception feature data is sent to the server-side control model, and the knowledge resource data and object service basic data are retrieved through the server-side control model. Through the server-side control model, service intent data, target object service profiles, and service personnel status profiles are generated based on the structured perception feature data, the knowledge resource data, and the object service basic data. Service decision data is generated based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. Based on the service decision data, a set of target intelligent agents is determined, auxiliary action data is generated through the set of target intelligent agents, and the auxiliary action data is sent to the interactive terminal; Collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

2. The object service assistance method based on multimodal interaction as described in claim 1, characterized in that, Acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through an edge-side multimodal model, including: Establish a multimodal acquisition session on the interactive terminal, and configure a session identifier, modality source identifier, acquisition time identifier, and terminal status identifier for the multimodal acquisition session; Based on the session identifier, multiple types of input segments are received, and according to the modality source identifier, the multiple types of input segments are divided into multiple modal segments among voice input segments, text input segments, image input segments and operation behavior segments, to generate multimodal interactive input data; By deploying a multimodal model on the interactive terminal, multiple types of end-side perception segments corresponding to the multiple modal segments are extracted from the multimodal interactive input data. The multiple types of end-side perception segments include multiple segments from speech semantic segments, text semantic segments, image object segments, and operation feature segments. Based on the acquisition time identifier and the modality source identifier, the multi-type end-side sensing segments are time-aligned and cross-modal associated using the end-side multimodal model, and acquisition status markers are added based on the terminal status identifier to generate cross-modal associated feature data; Through the edge multimodal model, service intent candidate information, target object pointing information, personnel status prompt information and scene trigger prompt information are extracted from the cross-modal association feature data, and collection status information is extracted from the collection status marker. The session identifier, modality source identifier, collection time identifier, service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information, and collection status information are encapsulated into structured perception feature data according to a field format that includes session field, modality field, time field, intent field, object field, personnel status field, scene trigger field, and collection status field.

3. The object service assistance method based on multimodal interaction as described in claim 1, characterized in that, The structured perception feature data is sent to the server-side control model, and the knowledge resource data and object service basic data are retrieved through the server-side control model, including: The structured perception feature data is encapsulated into edge-side transmission data, and the edge-side transmission data is sent to the server-side control model by the interactive terminal. The server-side control model parses the data sent from the terminal to obtain structured perception feature data, and extracts session identifier, target object pointing information, service intent candidate information, personnel status prompt information, scene trigger prompt information and collection status information from the structured perception feature data to generate service retrieval index data. The server-side control model generates knowledge retrieval index data and object service retrieval index data based on the service retrieval index data. Through the server-side control model, scenario knowledge data, organizational knowledge data, service resource data, and operation guidance data are retrieved from the knowledge resource storage area based on the knowledge retrieval index data, and knowledge resource data is generated based on the scenario knowledge data, the organizational knowledge data, the service resource data, and the operation guidance data; Through the server-side control model, the target object basic data, target object interaction record data, target object service status data, service personnel capability data, and service personnel task status data are retrieved from the object service data storage area based on the object service retrieval index data. Based on the target object basic data, the target object interaction record data, the target object service status data, the service personnel capability data, and the service personnel task status data, the target service basic data is generated. The knowledge resource data and the object service basic data are bound to the structured perception feature data according to the session identifier through the server control model.

4. The object service assistance method based on multimodal interaction as described in claim 3, characterized in that, Through the server-side control model, service intent data, target object service profiles, and service personnel status profiles are generated based on the structured perception feature data, the knowledge resource data, and the object service basic data, including: Through the server-side control model, the mutually bound structured perception feature data, knowledge resource data, and object service basic data are read according to the session identifier; Through the server-side control model, service intent candidate information, target object pointing information, personnel status prompt information, scene trigger prompt information, and collection status information are extracted from the structured perception feature data; Through the server-side control model, scene semantic data is generated based on the scene trigger prompt information and the knowledge resource data, and resource boundary data is generated based on the knowledge resource data; Service intent data is generated based on the service intent candidate information, the scene semantic data, the resource boundary data, and the collection status information through the server-side control model. Through the server-side control model, target object status data is extracted from the object service basic data based on the target object pointing information, and a target object service profile is generated based on the target object status data, the service intent data, and the knowledge resource data. Through the server-side control model, service personnel status data is extracted from the object service basic data based on the personnel status prompt information, and a service personnel status profile is generated based on the service personnel status data, the service intent data, and the knowledge resource data.

5. The object service assistance method based on multimodal interaction as described in claim 1, characterized in that, Through the server-side control model, service decision data is generated based on the service intent data, the target object service profile, and the service personnel status profile, including: Through the server-side control model, service target information, trigger scenario information and service urgency information are extracted from the service intent data; object activity status data, object demand status data and object service gap data are extracted from the target object service profile; and personnel capability status data, personnel task load data and personnel response preference data are extracted from the service personnel status profile. Service constraint data is generated based on the trigger scenario information, object demand status data, object service gap data, personnel capability status data, and personnel task load data through the server-side control model. Based on the service target information and the service constraint data, the server-side control model generates candidate data for service actions. The server-side control model generates service priority data based on the service urgency information, the object activity status data, and the personnel task load data. The server-side control model generates agent invocation conditions and auxiliary action output conditions based on the service action candidate data, service priority data, and personnel response preference data, and encapsulates the service action candidate data, service priority data, agent invocation conditions, and auxiliary action output conditions into service decision data.

6. The object service assistance method based on multimodal interaction as described in claim 1, characterized in that, Based on the service decision data, a set of target intelligent agents is determined; auxiliary action data is generated from the set of target intelligent agents; and the auxiliary action data is sent to the interactive terminal, including: Through the server-side control model, agent invocation conditions, service action candidate data, service priority data, and auxiliary action output conditions are extracted from the service decision data. Through the server-side control model, candidate agents are selected from the agent registration data based on the agent invocation conditions, and the action capabilities of the candidate agents are compared with the service action candidate data to generate candidate agent capability data. Based on the service priority data and the candidate agent capability data, the target agent set is determined from the candidate agents using the server-side control model. The server-side control model distributes the service action candidate data and the auxiliary action output conditions to the target intelligent agent set, and generates an action fragment set corresponding to the target intelligent agent set. The action fragment set includes any one or more types of fragments among task action fragments, resource action fragments, and interactive prompt fragments. Through the server-side control model, conflict resolution and presentation order arrangement are performed on the action fragment set based on the service priority data to generate auxiliary action data. The server-side control model converts the auxiliary action data into terminal display data and terminal interaction trigger data based on the auxiliary action output conditions, and then sends the terminal display data and terminal interaction trigger data to the interactive terminal.

7. The object service assistance method based on multimodal interaction as described in claim 1, characterized in that, Collect response feedback data generated in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile, including: The system collects terminal feedback content, object feedback content, and personnel feedback content in response to the auxiliary action data through an interactive terminal, and generates response feedback data based on the terminal feedback content, object feedback content, and personnel feedback content. The server-side control model extracts action response content, object feedback content, personnel feedback content, and action completion status from the response feedback data. Through the server-side control model, target object profile update data is generated based on the object feedback content, the action response content, and the action completion status. The target object service profile is then updated based on the target object profile update data to generate the updated target object service profile. The server-side control model generates updated service personnel profile data based on the personnel feedback content, the action response content, and the action completion status, and updates the service personnel status profile based on the updated service personnel profile data, thus generating an updated service personnel status profile. Through the server-side control model, subsequent object service requirement data is generated based on the updated target object service profile, subsequent personnel service status data is generated based on the updated service personnel status profile, and subsequent decision triggering data is generated based on the action completion status. The server-side control model generates subsequent service decision data for determining the set of subsequent target intelligent agents based on the subsequent object service demand data, the subsequent personnel service status data, and the subsequent decision triggering data.

8. An object service auxiliary device based on multimodal interaction, characterized in that, The object service auxiliary device based on multimodal interaction includes: The edge-side multimodal perception module is used to acquire multimodal interactive input data collected by the interactive terminal, and extract structured perception feature data from the multimodal interactive input data through the edge-side multimodal model; The server-side resource retrieval module is used to send the structured perception feature data to the server-side control model, and retrieve knowledge resource data and object service basic data through the server-side control model; The intent profile generation module is used to generate service intent data, target object service profiles, and service personnel status profiles based on the structured perception feature data, the knowledge resource data, and the object service basic data through the server-side control model. The service decision generation module is used to generate service decision data based on the service intent data, the target object service profile, and the service personnel status profile through the server-side control model. The intelligent agent action orchestration module is used to determine a target intelligent agent set based on the service decision data, generate auxiliary action data through the target intelligent agent set, and send the auxiliary action data to the interactive terminal; The feedback update module is used to collect response feedback data formed in response to the auxiliary action data, update the target object service profile and the service personnel status profile based on the response feedback data, and generate subsequent service decision data for determining the subsequent target intelligent agent set based on the updated target object service profile and the updated service personnel status profile.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multimodal interaction-based object service assistance program stored in the memory and executable on the processor. When executed by the processor, the multimodal interaction-based object service assistance program implements the steps of the multimodal interaction-based object service assistance method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores an object service auxiliary program based on multimodal interaction, which, when executed by a processor, implements the steps of the object service auxiliary method based on multimodal interaction as described in any one of claims 1-7.