Intention analysis and strategy generation method and device, equipment and medium

By identifying the interaction type and selecting the adaptive analytical model to process multimodal interaction data, generating text corpus and performing intention scoring analysis, the problem of inefficient multimodal interaction data processing in the existing technology is solved, and a fast and accurate customer demand list and interaction strategy generation is achieved, which improves service quality and response speed.

CN120492602APending Publication Date: 2025-08-15CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510589656.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing technology cannot efficiently identify and structure the multimodal interactive data, and lacks the ability to automatically generate customer intention scores and demand lists based on interactive content, which affects the personalized adaptation of insurance products and the efficiency of medical services.

Method used

By obtaining the metadata feature of the interactive data, identifying the interaction type, selecting the adaptive analytical model for analysis, generating text corpus, extracting user information from the text corpus and performing intent scoring analysis, generating a requirement list and interaction strategy.

Benefits of technology

It realizes the rapid and accurate generation of customer demand lists and interaction strategies, improves the efficiency and accuracy of customer information processing, reduces the workload of manual input and analysis, and improves customer service quality and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492602A_ABST
    Figure CN120492602A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semantic analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an intention analysis and strategy generation method, device, equipment and medium, and the method comprises the steps: obtaining interaction data, recognizing an interaction type according to metadata features, selecting a corresponding analysis model to analyze and process data, and generating a text corpus; current user information is extracted from the text corpus, user intention score analysis is executed based on an intention analysis strategy corresponding to the interaction type, and a current user intention score is generated; and generating a demand list according to the user information, the intention score and the interaction type, and generating and outputting an interaction strategy based on the user information and the demand list. Through interaction data processing and automatic intention scoring analysis, the demand list and the interaction strategy of the customer can be quickly and accurately generated, the customer information processing efficiency and accuracy are improved, the workload of manual input and analysis is reduced, and the customer service quality and the response speed are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic parsing technology, and in particular to an intent analysis and strategy generation method, apparatus, device and storage medium. Background Art

[0002] In the fintech business sector, insurance agents are an important bridge connecting insurance products and customers. In actual business processes, agents often need to communicate with customers through various means such as face-to-face visits, telephone calls, or social software (such as WeChat), and form preliminary communication records. These records come in various forms, including voice recordings, on-site photos, text notes, etc. In the existing process, agents need to manually organize the notes after the communication, extract basic customer information and fill it out in a standardized form or system input interface. Typical information includes customer name, age, contact information, ID number, and family structure. This process is not only highly repetitive and inefficient, but also prone to input errors or information omissions, seriously affecting work quality and customer service experience.

[0003] Furthermore, agents must rely on their own experience to subjectively judge the notes, analyze the customer's insurance intentions, and then complete a willingness questionnaire or needs assessment form. This model relies heavily on the agent's subjective ability, lacks unified evaluation standards, and makes it difficult to generate structured data for subsequent service processes. For example, different agents may have different understandings of the same voice content, resulting in inconsistent customer profiles and product recommendations, affecting the personalized adaptation of insurance products and the efficiency of subsequent service responses.

[0004] In the healthcare sector, service personnel face similar challenges. For example, when dealing with elderly clients, those with chronic illnesses, or multigenerational family members, customer needs are often complex, requiring the collection and processing of unstructured information from multiple rounds of communication, such as family medical history, previous insurance claims records, and health status descriptions. Currently, manual processing of this information and the determination of risk preferences and insurance needs is still largely reliant on manual labor. This operation is costly, and accuracy and timeliness are difficult to guarantee. This is particularly prone to processing bottlenecks in high-frequency business scenarios.

[0005] Overall, existing technologies have significant shortcomings in processing multimodal customer communication data, automatically extracting key information, and conducting structured management. They lack effective means to transform heterogeneous data such as voice, images, and text into standardized customer information. Furthermore, existing systems struggle to automatically identify potential customer needs and intentions based on communication content and complete standardized willingness assessments, which impacts subsequent service automation and the efficiency of business process collaboration. These issues urgently need to be addressed through intelligent technologies capable of cross-modal understanding and user intent modeling. Summary of the Invention

[0006] The main purpose of the present invention is to provide an intent analysis and strategy generation method, device, equipment and storage medium, aiming to solve the technical problems that the existing technology cannot efficiently identify and structuredly process multimodal interaction data, and lacks the ability to automatically generate customer intent scores and demand lists based on the interaction content.

[0007] To achieve the above objectives, the present invention provides an intention analysis and strategy generation method, comprising:

[0008] Acquire interaction data, and identify interaction types based on metadata features of the interaction data;

[0009] Selecting a corresponding parsing model based on the interaction type, and parsing the interaction data according to the parsing model to generate a text corpus;

[0010] Extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score;

[0011] generating a requirements list according to the current user information, the current user intention score, and the interaction type;

[0012] Generate an interaction strategy output based on the current user information and the requirement list.

[0013] Furthermore, to achieve the above-mentioned purpose, the present invention provides an intention analysis and strategy generation device, comprising:

[0014] An interaction identification module, configured to obtain interaction data and identify interaction types based on metadata features of the interaction data;

[0015] A multimodal parsing module, configured to select a corresponding parsing model based on the interaction type, and parse the interaction data according to the parsing model to generate a text corpus;

[0016] An intent scoring module is used to extract current user information from the text corpus and perform user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score;

[0017] A demand generation module, configured to generate a demand list based on the current user information, the current user intention score, and the interaction type;

[0018] A strategy output module is used to generate an interaction strategy output based on the current user information and the requirement list.

[0019] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an intent analysis and policy generation program stored in the memory and runnable on the processor. When the intent analysis and policy generation program is executed by the processor, the steps of the intent analysis and policy generation method described above are implemented.

[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an intent analysis and policy generation program is stored, and when the intent analysis and policy generation program is executed by a processor, the steps of the intent analysis and policy generation method as described above are implemented.

[0021] Beneficial effects: The present invention relates to the field of semantic parsing technology and can be applied to business scenarios such as financial technology and medical health. It discloses a method for intent analysis and strategy generation, including: obtaining interaction data and identifying the interaction type based on metadata features, selecting a corresponding parsing model to parse and process the data, and generating a text corpus; extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate the current user intent score; generating a requirements list based on user information, intent score and interaction type, and generating an interaction strategy output based on the user information and the requirements list. Through interaction data processing and automatic intent scoring analysis, the present invention can quickly and accurately generate a customer's requirements list and interaction strategy, improve the efficiency and accuracy of customer information processing, reduce the workload of manual input and analysis, and improve customer service quality and response speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0023] Figure 1 A schematic diagram of an application environment of the intent analysis and strategy generation method according to an embodiment of the present invention;

[0024] Figure 2 This is a flow chart of an embodiment of the method for intent analysis and strategy generation of the present invention;

[0025] Figure 3 Schematic diagram of functional modules of a preferred embodiment of the apparatus for intention analysis and strategy generation of the present invention;

[0026] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] The intention analysis and strategy generation method provided by the embodiment of the present invention can be applied in Figure 1 in an application environment, wherein the user terminal communicates with the server terminal through a network. The server terminal can obtain interaction data through the user terminal and identify the interaction type according to metadata features, select the corresponding parsing model to parse and process the data, and generate a text corpus; extract the current user information from the text corpus, and perform user intention scoring analysis based on the intention analysis strategy corresponding to the interaction type to generate the current user intention score; generate a requirement list based on the user information, intention score and interaction type, and generate an interaction strategy output based on the user information and the requirement list. The present invention can quickly and accurately generate a customer's requirement list and interaction strategy through interaction data processing and automatic intention scoring analysis, improve the efficiency and accuracy of customer information processing, reduce the workload of manual input and analysis, and improve customer service quality and response speed. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.

[0030] See also Figure 2 , Figure 2 This is a flowchart of an embodiment of the intention analysis and strategy generation method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than here.

[0031] like Figure 2 As shown, the intention analysis and strategy generation method proposed in the present invention includes the following steps:

[0032] S10, acquiring interaction data, and identifying the interaction type according to metadata features of the interaction data;

[0033] In this embodiment, interaction data includes, but is not limited to, text messages, audio, video, and accompanying behavioral trajectory data. In practice, interaction data typically originates from various customer communication channels, such as voice recordings from face-to-face interactions between offline agents and customers, call recordings from remote telephone conversations, text and image information transmitted via social platforms like WeChat or internal enterprise customer service systems, or multimodal video streams captured via video conferencing platforms. Each data source carries specific metadata features that serve as the foundation for identifying interaction types.

[0034] The metadata features of interactive data mainly include format type, communication protocol, content structure, encoding method, and the transmission context it carries. Audio data usually has features such as audio sampling rate, number of channels, and encoding type. Video data includes frame rate, resolution, audio and video synchronization identifiers, etc. Text data is represented by string format, character set identifier, message header information, etc. These metadata can be automatically extracted by parsing the data's transmission header, file header field, packet protocol information, or the context environment during the data reading process, and used for preliminary classification and judgment. For example, by identifying whether there is an audio waveform or channel identifier in the data frame, it can be quickly determined to be an audio interaction; if the encapsulation structure contains audio track and image frame information, it can be identified as a video interaction; if the data structure is a pure string and contains a natural language structure, it can be determined to be a text interaction.

[0035] Identifying interaction types is not only a data classification process but also a critical step in providing a basis for subsequent multimodal parsing model invocation. After type identification is complete, the corresponding parsing module is invoked to perform semantic content analysis. Therefore, this identification operation plays a decisive role in the correctness of the parsing chain. In practice, type identification can be achieved through rule-based matching methods, protocol field calibration logic, or initial training combined with shallow neural networks to achieve automatic classification of input data types.

[0036] In different scenarios, the method of acquiring interaction data can be technically adapted. For mobile applications, such as when insurance agents communicate with customers on-site through an app, the embedded system recording component can collect audio interaction data in real time, and automatically attach metadata such as the collection time, location, and device identification after the data recording is completed. In call center scenarios, the customer call process is connected to the system through the SIP protocol. The server automatically extracts the audio stream from the communication protocol and combines it with information such as call duration and channel ID to form a complete interaction data record. In the customer service IM system, all text interaction information is stored as a structured message object, including fields such as the sending time, sender ID, message content, and format identifier. It is uniformly extracted and initially formatted at the message receiving end.

[0037] To identify interaction types, a type parsing module based on preset rules can be used. For example, regularization rules can be used to determine whether the message body contains audio waveform encoding structures, video frame labels, or natural language paragraph markers. Alternatively, a lightweight neural network can be constructed to input interaction metadata vectors and output data type labels, thereby achieving structured type identification. In the case of mixed input of multiple data types, such as a video conversation containing both audio and image content, the primary interaction type can be prioritized based on data volume or main content proportion, and semantic diversion can be performed based on the primary type.

[0038] It can also be adapted and optimized for special environments. For example, in financial data access systems, by connecting with the customer relationship management (CRM) system, while collecting interaction data, it can also associate the current session with the corresponding customer historical information and service nodes to supplement the metadata context; in medical scenarios, contextual information from electronic health records (EHR) can be embedded to enhance the subsequent semantic parsing module's ability to recognize professional terms.

[0039] Example: In the healthcare field, when an agent explains a health insurance plan to a customer via video call, the system collects the customer's voice inquiries, screen annotation behaviors, and video images in real time. The video stream contains frame rate and synchronization timestamp information, and the audio track carries sampling rate and compression format identifiers. The system automatically identifies it as a video interaction type and calls a joint parsing model including OCR and speech recognition for multimodal semantic extraction.

[0040] In the financial field, when customers inquire about pension products by phone, the system collects the call audio and identifies the audio type based on the communication protocol information. It completes metadata calibration through the audio format characteristics and message time nodes, identifies it as a voice interaction type, and switches to the voice recognition parsing path for transcription processing, providing an input basis for subsequent intent analysis and risk preference modeling.

[0041] By automatically identifying interaction types based on metadata features of interaction data, not only can data input be automatically classified, eliminating manual preprocessing, but it also provides structural support for precise matching of subsequent multimodal parsing paths. This enables the system to adapt to customer interaction content from multiple channels and formats, improving the adaptability and versatility of input data. The pre-parsing identification mechanism established based on this process can significantly improve system processing efficiency and reduce the risk of parsing failure or information loss due to type misjudgment in the data processing chain.

[0042] S20, selecting a corresponding parsing model based on the interaction type, and parsing the interaction data according to the parsing model to generate a text corpus;

[0043] In this embodiment, the interaction type refers to the interaction form category to which the current interaction data belongs, which generally includes three basic forms: voice conversation, video conversation, and text conversation. The original input of a voice conversation is audio data, a video conversation contains image frames, audio tracks, and possible real-time interaction elements, and a text conversation generally refers to text messages input through instant messaging tools, forms, or web pages. The identification of the interaction type is usually completed based on the metadata features in the interaction data, such as the encoding format of the audio file (such as MP3, WAV), the video container format (such as MP4, MKV), or the encoding tag of the text data (such as UTF-8, HTML paragraph structure, etc.). It can also be judged by the context attributes of the data entry, such as the media type declaration in the request header, the identifier of the input channel, etc.

[0044] After identifying the specific interaction type, it is necessary to match and select an adaptive parsing model based on the interaction type. The types of parsing models include but are not limited to pre-trained speech recognition models, video content understanding models, text normalization and segmentation models, etc. The parsing model has the characteristics of adapting the input port to the modality of the interaction data and matching the processing logic with the semantic extraction capability. For example, for speech-type interaction data, a speech recognition model is preferred, which contains a speech signal processing submodule, a phoneme recognition module, and a contextual semantic correction module; for video-type data, a graphic-text fusion model or a visual-to-text model with OCR capabilities is usually selected; and text-type data is more suitable for an NLP text segmentation and parsing model that includes encoding standard conversion and sentence structure segmentation functions.

[0045] After inputting the interaction data into the corresponding parsing model, the model's internal recognition mechanism converts audio, images, or encoded text content into a processable standardized text corpus. Text corpus is a uniformly structured textual expression that serves as the fundamental input structure for subsequent intent analysis and strategy generation. It typically includes the text of the conversation, annotated time information, sentence classification labels, and entity extraction results. This process is not only a key step in normalizing multimodal input but also essential for ensuring the accuracy and traceability of subsequent intent recognition.

[0046] The technical implementation of text corpus generation requires ensuring structural consistency and timestamp alignment. For example, in video and voice data, the model must segment and sequence multiple speech segments. In text data, encoding formats must be corrected, content segmented into standard segments, and semantic clarity must be enhanced through methods such as word density and grammatical structure recognition.

[0047] Different parsing paths can be set for different types of interaction data. For example, when the interaction data is detected as speech, the system can call a speech recognition model based on a deep neural network, such as a speech model based on a Transformer structure. After passing through a preprocessing module (such as silent segment cropping and waveform noise reduction processing), it is input into the speech recognition backbone model, outputs word-level transcription results, and then combines the context language model for text correction, ultimately generating a unified structure of speech and text corpus.

[0048] If the interaction type is video, the session processing logic must first extract keyframe images from the video and use the OCR model to detect and recognize the text content therein; it must also simultaneously extract the voice track in the video, transcribe it into text through the voice recognition module, and then use the timestamp alignment algorithm to merge the image text and voice text along the timeline to form a video text corpus with strong temporal consistency.

[0049] For text conversation types, first identify its original encoding format, use the encoding conversion module to perform unified standard processing (such as unified conversion to UTF-8 format), and then use the sentence segmentation model or regular expression model to perform structured segmentation on the text to construct a structured text corpus with sentence patterns, logical paragraphs, and entity recognition tags for use by subsequent modules.

[0050] In different implementation environments, the structure and processing methods of the parsing model can be further optimized based on specific business scenarios. For example, in scenarios where device computing power is limited, a lightweight speech recognition model or an edge OCR model can be used instead of a cloud-based model. In scenarios where text timeliness is a high priority, common semantic templates can be preloaded to improve parsing efficiency.

[0051] Example: In the healthcare field, a user communicates remotely with a doctor via video chat, checking health declaration items and providing verbal instructions. The system first identifies the video data type, extracts the options clicked by the user and the content of the speech, and generates a clear video parsed text corpus, which is then used to determine whether the user meets the conditions for exemption from certain medical examinations.

[0052] In the financial field, an insurance agent communicates with a customer over the phone, and the customer asks by voice, "What materials are needed to claim for a major illness?" The system recognizes it as a voice type, extracts the content, transcribes it into text corpus, and automatically labels it as a consulting request for subsequent intent scoring and response strategy generation, thereby improving the degree of process automation.

[0053] In the enterprise marketing assistance scenario, the system receives user input in text form, "I want to cancel my insurance plan," and identifies it as text interaction data. After unified coding and structured segmentation, it classifies the text corpus as a cancellation operation with a high level of intent, providing precise support for the subsequent generation of retention plans or customer churn warning strategies.

[0054] By strictly mapping interaction types to parsing models and executing a structurally consistent, semantically aligned parsing process on multimodal interaction data, we can effectively transform raw unstructured interaction data into standardized text corpora. This conversion process reduces manual collation costs and the accumulated errors in multimodal data conversion, enabling subsequent modules such as intent recognition, demand extraction, and strategy generation to achieve efficient and accurate semantic processing within a unified data structure. This not only improves the processing efficiency of customer interaction data, but also demonstrates excellent scalability and stability in the fusion of multi-source heterogeneous data.

[0055] S30, extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score;

[0056] In this embodiment, extracting current user information from text corpus involves identifying semantic clues related to user identity, behavioral characteristics and preference tendencies in the standardized language data. This information may include multiple dimensions such as occupation, age, family structure, regional characteristics, historical behavior records, and expressions. The source of text corpus is not limited to a certain type of interaction method. It can be oral expression content obtained through voice transcription, or it can be text conversation, graphic data analysis results, or behavior records recognized in video scenes. Since these data are presented in various forms and have flexible semantic expressions, they usually do not have structured labels. Therefore, information extraction cannot rely on fixed fields, keywords, or template sentences, but should be flexibly judged based on contextual logic and semantic flow. In the processing process, technical means such as lexical analysis, entity recognition, hierarchical relationship recognition, and text semantic modeling are often combined to improve the accuracy and completeness of user information extraction.

[0057] User information extraction can be implemented in conjunction with a deep semantic understanding module. For example, named entity recognition based on a language model can be used to annotate occupations, organization names, or role titles. Dependency parsing can then be used to extract structural relationships. For example, the phrase "I work in a hospital" can identify occupational information, while "We are a family of five" can infer family size. In some sentences, such as "I recently took my child for a checkup" or "I plan to address my retirement security issues," while family structure or occupational status is not explicitly stated, contextual clues can be used to semantically restore implicit features for subsequent behavior modeling and demand prediction.

[0058] On the basis of obtaining user information, it is necessary to select an appropriate intent analysis strategy according to different interaction types. The intent analysis process is not a universal processing path, but rather selects analysis methods based on the expression method, input form, and interaction context under different interaction channels. For example, in voice conversations, the model may focus on identifying speech speed, pitch fluctuations, emotional excitement, and the proportion of interrogative sentences; in text and video, the system pays more attention to user operation behavior, screen markings, and visual focus areas; and in text conversations, it makes judgments based on dimensions such as keyword distribution, sentence structure, and emotional polarity. These strategies have logical differences, and are associated with different model architectures and scoring formulas. The evaluation indicators also have different emphases due to different interaction types.

[0059] User intent scores aren't generated solely through a single scoring logic; they typically combine multiple perceptual indicators, behavioral signals, and semantic units. For example, in the phrase "I urgently need this insurance product," the system considers not only the keyword strength of "urgently," but also the speed and degree of voice rise during expression to comprehensively reflect the intensity and urgency of the emotion. For scenarios where a user circles a specific clause on the screen and adds the annotation "Please explain clearly," a specific intent weight is generated based on the highlighted area, screen position, time, and text content.

[0060] The resulting current user intent score can be presented as a multidimensional evaluation vector or mapped to a priority level, intent classification label, or policy trigger set. This derivative form seamlessly integrates with subsequent demand analysis and interaction strategy matching modules, forming a continuous, closed-loop service response path. The entire scoring system supports on-demand adjustment of evaluation dimensions and model structure to meet the accuracy and real-time requirements of different business processes.

[0061] In one specific implementation, a joint model system can be used to complete user information extraction and intent scoring. The current user information extraction module can integrate an entity recognition component based on context-aware language models such as BERT, perform sequence annotation after completing word embedding, output a candidate list of fields such as occupation and age, and confirm the field value in conjunction with a rule enhancement module. A rule-based text inference module can also be integrated to identify the semantic intent behind expressions such as "just retired" and "the three of us", thereby achieving structured expression of implicit information.

[0062] In the intent analysis strategy module, the system can route processing paths according to the interaction type. For example, in a voice interaction scenario, the system uses an acoustic model to extract pitch changes, pause duration, and stress position, and combines the proportion of interrogative sentences and imperative sentence structure in the transcribed text to construct an intent factor, and then outputs the voice intent score through the training model. In video scenarios, the OCR module can be used to extract text content, and the annotation area can be located in combination with coordinate information. The focus object and emotional orientation can be determined through the image-text alignment model, and the intent strength can be calculated using rule matching or neural networks. In text conversations, keyword libraries and grammatical analysis tools can be introduced to quantify dimensions such as request categories, verb intent, and semantic emotions in messages to form a structured intent score.

[0063] This scoring output module can further support a dynamic policy update mechanism, allowing parameter weights to be adjusted through real-time user feedback or model iteration to enhance adaptability. Technical capabilities such as multilingual input, terminology standardization, and typo tolerance can also be embedded into the implementation path to improve the model's stability and applicability in real-world environments.

[0064] Example: In a healthcare scenario, a patient might communicate with a health advisor via voice, saying, "I've been feeling dizzy lately, and I'm a little worried it might be a neurological problem." The system can extract words like "dizzy," "worried," and "nervous" from the voice. Combining these with changes in the voice's tone and the user's occupation ("teacher"), the system can generate a high intent score favoring neurological examinations and chronic disease consultations, and assist in recommending relevant insurance coverage.

[0065] In a financial scenario, a customer views a terms page through a video demonstration platform, highlights the "disclaimer" on the screen, and then verbally says, "I don't quite understand this." The system can recognize their annotation behavior and semantic questions, and combine the age field "55 years old" and the family structure "retired" to determine that they have a clear concern about the scope of protection, output a high-intensity confirmation intention, and provide guidance for the next step of strategy generation.

[0066] By building adaptive parsing and scoring logic for multi-type interaction data, we can accurately extract key identity and preference information from user expressions and, combined with corresponding intent analysis strategies, calculate a dynamic intent strength index. This scoring approach, based on differentiated interaction type processing, not only improves the model's understanding accuracy but also enhances its understanding of user context, implicit needs, and behavioral intent. This provides a reliable foundation for subsequently generating highly compatible service responses or insurance recommendations, significantly reduces manual judgment errors and missed judgments, and effectively improves overall automation.

[0067] S40, generating a requirements list according to the current user information, the current user intention score and the interaction type;

[0068] In this embodiment, the process of generating a list of needs involves three key input information: current user information, current user intent score, and interaction type. Among them, current user information refers to data fields related to user identity, background, and behavioral tendencies, such as occupation, age group, family structure, and regional preferences, identified by parsing the expression content in the text corpus. This information is not directly extracted from fixed fields, but is the result of contextual recognition based on semantic understanding, syntactic relations, or implicit expressions. For example, "I just retired" may infer that the age group is over 60 years old, and "I care about the health protection of my two children" can deduce family structure and protection preferences.

[0069] The current user intent score is a scalar measure of intent strength, or a structured set of hierarchical labels, generated by the system after completing multi-dimensional intent recognition analysis. It quantifies the current user's potential demand activity, priority, or urgency of expression within a specific time period. This score can be derived from a comprehensive analysis of dimensions such as the intensity of expression in text, the density of checkboxes in video interactions, and the emotional fluctuations in voice intonation. It is typically calculated using a set of scoring models or a weighted fusion mechanism.

[0070] Interaction type is the classification of interaction channels determined by the system after metadata recognition. It is typically categorized into voice, video, and text conversations. This classification plays a key role in the subsequent structure of requirement templates, priority hierarchy logic, and policy engine interface selection. Each interaction type corresponds to significantly different user behavior characteristics, information density, and operation methods, so the requirements list construction strategy should be adapted accordingly.

[0071] When generating a needs list, we first need to select a matching needs category structure based on the interaction type. Different interaction modes correspond to different user expression styles and interaction depths. For example, voice interactions tend to be more consultative, video interactions are more about confirming contract or product terms, and text interactions typically carry a strong intention of providing operational instructions. Next, we combine contextual elements from the current user information. For example, the occupation field can map specific industry-specific protection items, the age field can trigger lifecycle stage requirements, and the family structure field can link to child, elderly, or spouse protection modules. This will construct a preliminary set of needs items centered on user characteristics.

[0072] Furthermore, the current user intent score will sort or prioritize the matched requirements. The system can set different priority categories based on the score distribution range. For example, intent scores above a set threshold can be marked as "urgent," those in the medium range as "high," and those below the standard as "regular." This ultimately generates a structured requirement list entity, including fields such as requirement category labels, user-related items, and priority indicators to support the subsequent invocation of the policy output module.

[0073] In practice, a mapping table between interaction types and requirement categories can be constructed to quickly categorize interaction scenarios. For example, voice interactions can be mapped to "consultation requirements," video interactions to "confirmation requirements," and text interactions to "operation requirements." This rule can be loaded into the classification engine during the initial service phase. The system can also dynamically adjust classification strategies through the configuration management platform to adapt to the changing needs of different enterprise business processes.

[0074] To map user information to demand items, we can build a demand item knowledge base indexed by fields such as occupation, family structure, and age, and use conditional rules or machine learning models for correlation and matching. When the occupation field is "teacher," education fund savings requirements can be activated; when the family structure is "multiple children," product items such as child health insurance and school insurance can be triggered. These rules can be manually configured by business personnel or automatically generated through statistical modeling based on historical user data.

[0075] Prioritization of intent scores can be mapped using threshold intervals. For example, scores greater than 1.5 are marked as "urgent," scores between 1.0 and 1.5 as "high," and scores below 1.0 as "normal." Percentile ranking can also be used to dynamically segment intent distribution to accommodate overall fluctuations in user behavior across different data sets or business cycles. The structured encapsulation of the requirements list can be output in standard JSON format, with field names, output format, and field combinations configurable based on subsequent user requirements.

[0076] Example: In the healthcare sector, an agent uploads a voice recording after communicating with a user. The transcribed text contains expressions such as "I've been running to the hospital lately, and my child's checkup had a problem." The system identifies the related expressions "children," "checkup," and "hospital" through context, inferring that the family structure includes children, and constructs a list of needs centered around "children's medical insurance" and "major disease screening." The intent score is assigned a high priority tag due to the inclusion of keywords such as "recently" and "problem," and the resulting list of needs prioritizes relevant protection solutions.

[0077] In a financial scenario, a user signed a contract remotely with an agent via video. In the video sharing interface, they circled "Revenue Calculation Model" and handwritten a comment, "Please explain this section in detail." After identifying the selection and comment, the system determined that the user had a misunderstanding of the product's revenue structure. Based on the user's occupational information ("Business Owner"), it added "High Net Worth Investment Planning" and "Product Structure Analysis" to the list of confirmation requirements, assigning a medium-high intent score to guide the subsequent personalized financial advisory service matching process.

[0078] By introducing a structured requirements list generated from three types of information: interaction type, user information, and intent score, we can proactively analyze and prioritize user needs. This eliminates the traditional manual re-entry of note information and label assignment, reduces tedious operations and subjective bias, and ensures semantic coherence between the requirements and the context in which they are expressed. This structured output enables subsequent modules to perform strategy recommendations, product matching, or task assignment based on a unified interface, significantly improving the automation level and responsiveness of the overall process.

[0079] S50: Generate an interaction strategy output based on the current user information and the requirement list.

[0080] In this embodiment, the generation logic of the interaction strategy output relies on the comprehensive judgment and strategy matching process of the current user information and the structured fields in the demand list. The current user information usually contains static attributes (such as occupation, age group, geographical location, family structure) and dynamic attributes (such as the time of the last interaction, preference tags, service history, etc.). This information constitutes the core foundation of the user portrait. The demand list comes from the analysis output of the user's expressed intention in the previous link. It already has structured fields such as demand category, item, priority, etc., and has a high degree of semantic aggregation and goal orientation.

[0081] Generating an interaction strategy isn't simply a matter of applying a single strategy template. Instead, it involves dynamically generating one or more strategy combinations, taking into account the compatibility between user attributes and current needs. The structure of the strategy output can include recommended product combinations, content prompts, guidance node configuration, content push paths, page component ordering, and even the branching structure of service processes. Different fields have multidimensional mapping paths, such as occupational information driving the selection of a script template, age structure influencing risk level prompts, and family structure influencing the order of service recommendations.

[0082] The strategy generation process involves rule matching, weight evaluation, context fusion, priority filtering, and logical sorting of combinations of multiple fields. The fields in the current user information are not only used as entry points for condition matching, but are also used to constrain the effectiveness of the recommendation strategy after strategy generation. For example, if the current user is a medical practitioner, recommendations may need to avoid ordinary health insurance products and focus on customized plans with high deductibles. If the user is over 60 years old, the system should incorporate risk warning mechanisms and underwriting guidance into the strategy to avoid misleading insurance applications.

[0083] Policy output can also dynamically adjust response levels based on the priorities set in the request list. High-priority requests can trigger direct human intervention, medium-priority requests can be matched with regular content push, and regular-priority requests can be queued for processing by the automated task system. The system should also consider resource constraints and concurrent task status when generating policies. Some policies can use delayed triggering or be dynamically adjusted based on subsequent user feedback.

[0084] Policy output must not only be used for internal system calls but also support multi-port, multi-modal, and multi-role scenarios. The system must design a policy output framework that can adapt to various downstream usage methods, such as mobile terminal display, manual customer service interfaces, push engines, or process robots, to ensure the coherence of interactive links and the consistency of response paths.

[0085] In one implementation, the policy generation engine constructs a mapping tree between user information fields and fields in the wish list, then uses a configured decision table or rule engine to execute conditional judgments. For example, if the user information fields include "Occupation = Teacher," "Age = 35," and "Family Structure = Two Children," and the wish list includes "Education Fund Insurance" with an "Urgent" priority, the system generates the following policy combination: push a product information card labeled "Children's Education Fund Insurance," initiate a conversational call to the education reserve planning service, invoke the corresponding product follow-up process in the CRM system, and generate structured recommendations for the agent.

[0086] A progressive policy tree can also be generated based on a multi-round dialogue environment. During the initial contact, only high-priority policies are output. After the user completes the confirmation operation or supplements the information, the policy node content is updated based on the new input, achieving dynamic policy evolution.

[0087] It is also possible to build a strategy recommendation model based on historical successful strategy cases, and use models based on collaborative filtering or graph neural networks to achieve structural matching between user similarity portraits and strategy paths, thereby recommending high-success rate strategy output solutions under similar user paths.

[0088] The policy output format supports JSON structure or policy object format. The content fields include policy objectives, reach channels, execution logic, recommended paths, trigger conditions and expiration mechanisms, etc., ensuring that each policy result is explainable, executable and traceable.

[0089] Example: In a healthcare scenario, a user uploads a voice utterance, which is then transcribed and generated into text. The system identifies the user as a 40-year-old female with two children, who recently requested information about "physical checkup packages" and "tumor marker screening." The list of needs includes "basic health checkup" and "screening for common diseases in middle-aged and elderly women," with a "high" priority. During the policy generation phase, the system uses this information to generate the following policy: The "Customized Physical Checkup Package for Women" will be displayed on the push recommendation page, breast cancer screening content will be added to the configuration policy, and a limited-time appointment guide button will be added. If the user does not click on the button for more than 90 seconds, a customer service prompt will be displayed.

[0090] In a financial business scenario, a user filled out a text-based policy transfer application, adding a note stating, "I need to optimize my financial plans and don't want to waste my old policy." The system identified the user as single and self-employed. Based on the intent score, it inferred a strong need for asset security and liquidity. The system then developed the following recommendations in its policy output: a "recommended portfolio of value-preserving annuity insurance that can be surrendered and transferred." The recommendation plan highlighted the "locked-in return + low-threshold transfer" clause, and the notification policy was set to "send once before 12:00 PM on weekdays, and repeat the next morning if unread." This policy link was automatically linked to the customer follow-up plan module in the CRM system.

[0091] By introducing an interaction strategy generation mechanism based on current user information and a structured requirements list, we achieve a fully closed loop from user intent identification to personalized response strategy design. This significantly reduces manual judgment costs and response delays, avoiding the service disconnection caused by traditional methods due to information fragmentation and unified response models. The structured strategy results not only improve interaction response efficiency but can also be used as a data asset for subsequent model training and business analysis, significantly enhancing customer service automation and conversion rates.

[0092] The present invention relates to the field of semantic parsing technology and can be applied to business scenarios such as financial technology and medical health. It discloses a method for intent analysis and strategy generation, including: obtaining interaction data and identifying the interaction type based on metadata features, selecting a corresponding parsing model to parse and process the data, and generating a text corpus; extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate the current user intent score; generating a requirements list based on user information, intent score, and interaction type, and generating an interaction strategy output based on the user information and the requirements list. Through interaction data processing and automatic intent scoring analysis, the present invention can quickly and accurately generate a customer's requirements list and interaction strategy, improve the efficiency and accuracy of customer information processing, reduce the workload of manual input and analysis, and improve customer service quality and response speed.

[0093] In one embodiment, the above step S20 includes:

[0094] S201: When the interaction type is a voice conversation, perform noise reduction processing on audio data in the interaction data, extract a track feature vector from the noise-reduced audio data, perform phoneme alignment and context error correction on the track feature vector using a pre-trained speech recognition model, and generate a first speech transcription text;

[0095] S202, when the interaction type is a video conversation, extracting key frame images from the video data in the interaction data, identifying text content in the key frame images using a text recognition model, performing voice transcription on the audio track of the video data to generate a second voice transcription text, and aligning and fusing the text content with the second voice transcription text along a time axis to generate a video parsed text;

[0096] S203, when the interaction type is a text conversation, identifying the encoding format of the original text in the interaction data, decoding and converting the original text based on the encoding format to generate text data in a unified encoding format, and segmenting the text data in the unified encoding format to generate structured conversation text;

[0097] S204: Perform cross-modal semantic alignment processing on the first speech transcription text, the video parsed text, and / or the structured conversation text to generate a standardized text corpus.

[0098] In this embodiment, after acquiring interaction data, the interaction type is first identified based on the metadata features carried in the interaction data. This identification can refer to multiple sources, including but not limited to protocol header fields, encapsulation format, media encoding identifier, duration threshold, frame rate parameters, channel type, and other sources. This determination does not rely on user subjective declarations; instead, the system automatically performs type mapping through analysis, ensuring consistent and accurate distinction between voice, video, and text conversations.

[0099] After the interaction type is identified, the corresponding parsing model is selected based on the type. The parsing model selection mechanism can pre-set a model-ID mapping table or dynamically call a policy engine to determine the optimal model path. The parsing model is a computational framework that performs semantic transformation on the information in the interaction data. Its essential function is to convert the original multimodal signal into a standardized corpus resource that can participate in subsequent semantic calculations. Different types of interaction data vary significantly in structure, expression dimensions, and carrier form, and therefore require matching different parsing models for processing.

[0100] When the interaction type is a voice conversation, the parsing model performs noise reduction processing on the original audio data to improve the subsequent recognition accuracy. Noise reduction methods may include spectral subtraction, beamforming, deep learning acoustic model residual filtering and other technical paths. After the audio data is denoised, the audio track feature vector is extracted. This vector can be a Mel-frequency cepstral coefficient, a short-time Fourier transform coefficient, or a feature embedding vector output by the end-to-end model. The feature vector is processed by the pre-trained speech recognition model to complete phoneme alignment and contextual error correction. The error correction link can combine the language model structure to repair contextual ambiguity, common confusing words, or phrase-level continuous misrecognition, and finally output the first speech transcription text. The text retains the original speech content but has stronger computability.

[0101] When the interaction type is a video conversation, keyframe images must be extracted from the video data. The keyframe extraction strategy can be based on criteria such as still frame density thresholds, image histogram changes, and motion region stability. A text recognition model is called in each frame to identify the text within the image and output the text-to-image expression field. Simultaneously, the audio track of the video file is transcribed to generate a second speech transcript. The text-to-image recognition results and the speech transcript are aligned and fused based on timestamps to construct a video parsed text. This fusion process not only matches time points but also considers semantic consistency and cross-modal directivity. For example, a user's annotation in the video, "Please explain this," is synchronized with the speech content at that moment to establish the context of the expression.

[0102] When the interaction type is a text conversation, the encoding format of the original text is first identified. Multiple text encodings may exist, such as UTF-8, GB2312, and ISO-8859-1. Encoding differences exist when transmitting across different platforms, necessitating unified conversion. After identification, decoding and conversion processing is performed based on the format to generate text data in a unified encoding format, ensuring no garbled characters or formatting offsets. The text is then semantically segmented to generate structured conversational text. Segmentation is not limited to punctuation and length; auxiliary mechanisms such as command intent recognition, turn-taking recognition, and contextual topic demarcation can also be incorporated to form stable semantic units for downstream analysis.

[0103] Speech transcription, video parsing, and structured conversational text each have their own unique characteristics in terms of semantic structure, linguistic style, and contextual coherence. Therefore, further cross-modal semantic alignment is necessary. This alignment goes beyond simple temporal overlap and should also consider semantic complementarity and behavioral consistency. For example, when there is a shift or overlap in meaning between speech expression and video annotation, a language model or graph neural network architecture is needed to determine the information fusion path. After alignment, a standardized text corpus is generated with a unified semantic expression structure that supports operations such as tagging, slot filling, and logical matching. This serves as the standard input for subsequent modules such as user information extraction, intent analysis, and policy recommendation.

[0104] In one specific implementation, the parsing model can be called using preload, on-demand loading, or streaming methods. For low-resource deployment on the client side, a lightweight ONNX model structure can be used. For high-concurrency server-side tasks, a multi-threaded asynchronous inference engine can be deployed. Audio feature extraction can use MFCC or Wav2Vec2.0 feature representations, and image OCR can be implemented based on a deep convolutional network or a visual Transformer model. Keyframe image extraction can dynamically adjust the frame rate using interval sampling, content change detection, or action recognition models. When integrating speech transcription with images, the alignment strategy can use an alignment algorithm based on longest common subsequence matching (LCSM) or train an end-to-end alignment network model to improve the stability of cross-modal integration. In medical scenarios, the segmentation strategy for structured conversational text can incorporate specialized medical terminology dictionary matching. In financial scenarios, it can be combined with an insurance policy terminology recognition model to automatically tag key fields such as "liability exemption" and "insurance amount" and organize the content structure by business dimensions. The cross-modal alignment operation can integrate the multimodal Transformer model and capture the correspondence between the semantics of images and text through the attention mechanism; the output of the standardized text corpus can define the field structure in JSON Schema, including semantic units, modal sources, time anchor points and other information, to ensure seamless data integration with downstream modules.

[0105] Example description: In a healthcare scenario, when a user conducts a remote video consultation with a health management consultant, the system identifies the title of the "Informed Consent Form for Gastrointestinal Endoscopy" and the checkbox "Agree to Painless Operation" from the keyframe images extracted from the video. At the same time, the user's statement "I want to ask if painless operation is more expensive" is obtained through audio transcription, and this expression is transcribed into the first voice transcription text. The two types of data sources overlap on the timeline. The semantic alignment model is used to associate the check operation in the image with the question content in the voice, and it is identified that the user's actual focus is "the cost difference of painless gastrointestinal endoscopy." In the structured corpus output, such expressions are uniformly converted into standard semantic units "Users raise questions about the cost of painless gastrointestinal endoscopy", with semantic source labels (image, voice), modal confidence scores and time anchors, which are used to drive the subsequent generation of explanation suggestions and package recommendation engine calls.

[0106] In financial services scenarios, insurance agents engage in multimodal communication with customers, including phone explanations, WeChat text and image messaging, and remote video demonstrations. The system captures the customer's voice transcription, "Can each of my two children be covered by this education insurance policy?" and converts it into a first-line audio transcript. A WeChat screenshot, through optical character recognition (OCR), reveals the product terms and conditions, including the statement "Multiple children can be covered separately within the same household." In the shared video, the user underlines the phrase "Separate share per person" and annotates "Confirm this section." These three data sources, semantically consistent and temporally continuous, are fused into a single semantic entry, "User confirms that multiple children can be covered separately," through a cross-modal semantic alignment mechanism. This data is then formatted into standardized text for use by the customer demand tag recognition module, enabling subsequent customer profiling optimization and strategic content compilation.

[0107] This embodiment, by building a unified parsing model system covering multiple interaction types such as voice, video, and text, can stably generate standardized text corpora for further semantic computation, even in contexts where data source input is uncertain and content presentation is complex and frequently changing. This text corpus not only achieves structured transformation at the content level but also builds a temporal semantic bridge between modalities, enabling subsequent modules such as user intent analysis, information extraction, and strategy generation to establish a common processing benchmark across different modalities, significantly reducing the overall system complexity and the risk of error transmission.

[0108] In one embodiment, the above step S30 includes:

[0109] S301, when the interaction type is a voice conversation, extracting interrogative sentences and imperative sentences from the text corpus as demand description fields;

[0110] S302, analyzing the interrogative word density and tone intensity value in the requirement description field to generate a speech intent score;

[0111] S303, determining a user attribute weighting factor based on the occupation, age, and / or family structure in the current user information, and multiplying the voice intent score by the user attribute weighting factor to generate a current user intent score;

[0112] S304, when the interaction type is a video conversation, identifying a check box and handwritten annotations in a screen sharing area in the text corpus as a requirement confirmation field;

[0113] S305, counting the frequency of annotation selections and the relevance of annotations in the requirement confirmation field to generate a video intent score;

[0114] S306, determining a user attribute weighting factor based on the occupation, age and / or family structure in the current user information;

[0115] S307, multiplying the video intent score by the user attribute weighting factor to generate a current user intent score;

[0116] S308, when the interaction type is a text conversation, classifying the text corpus into message types, and marking a demand urgency field in the text corpus according to the message type;

[0117] S309, counting the keyword matching degree and sentence complexity in the demand urgency field to generate a text intent score;

[0118] S310 , determining a user attribute weighting factor based on the occupation, age, and / or family structure in the current user information, and multiplying the text intent score by the user attribute weighting factor to generate a current user intent score.

[0119] In this embodiment, the current user information is extracted based on the text corpus, and the intent analysis process is performed according to the content structure and expression characteristics of different interaction types, and finally a quantifiable user intent score is generated, which is used as a logical precondition for subsequent demand modeling and strategy output. The current user information includes but is not limited to attributes reflecting the user's identity background, such as occupation, age, and family structure. This information can be obtained from the text corpus through named entity recognition, semantic reasoning, or syntactic parsing. The text corpus itself may be composed of speech recognition, video transcription, or original structured text. Therefore, the challenges of language non-structurality and modal intersection need to be taken into account during the parsing process.

[0120] When the interaction type is a voice conversation, it is often characterized by irregular spoken expressions and frequent changes in tone. Therefore, interrogative and imperative sentences are preferentially extracted as demand description fields. These sentences often appear in contexts expressing motivation, requests, or dissatisfaction and have a high user intent density. The interrogative word density can be calculated by identifying the ratio of the frequency of interrogative expressions (such as "what," "whether," and "how") in the corpus relative to the total text length. The tone intensity value can be combined with the occurrence of question-ending words and emotional words (such as "must" and "hurry") to form a quantitative indicator for weighting. The voice intent score is a weighted combination of the above two indicators, which is then multiplied by the weighting factor mapped from the current user information to obtain the user intent score.

[0121] When the interaction type is a video conversation, because the information expressed includes both visual markers and verbal content, it is important to focus on analyzing interactive behaviors within the screen sharing area. Checkboxes and handwritten annotations are explicit forms of user intent input, and their content and coordinates can be located through image recognition technology. The frequency of annotation checkboxes reflects the user's active participation, while annotation relevance calculates the semantic relevance between the annotation content and the background terms to determine whether it focuses on a specific topic. The video intent score is calculated by combining these two indicators. The score is then modified by a multiplicative factor based on the occupation, age, and family structure parameters of the current user information to produce comparable intent results.

[0122] When the interaction type is a text conversation, the first thing to do is to classify the message types because the language structure is relatively complete. Common categories include questions, statements, and instructions. The message type determines how the urgency of the demand is marked, which in turn guides the downstream intent modeling strategy. The urgency field quantifies the intensity of intent expression through keyword recognition (such as "immediately" and "as soon as possible") and sentence structure complexity (such as syntactic nesting and modification depth). The text intent score is calculated based on keyword matching and sentence complexity, and then multiplied by the attribute factor extracted from the current user information to generate the current user intent score.

[0123] In voice interaction mode, transcribed text can be obtained by accessing a pre-trained speech recognition engine. This is then combined with a dependency parsing model to identify sentence types and semantic roles, thereby extracting interrogative expressions. Question word density analysis can be performed based on a predefined vocabulary or BERT embedding similarity matching. Tone intensity assessment can be integrated with sentiment changes across multiple conversations to predict the intensity of the conversation. In the intent score output stage, a neural network regression model is used to learn the multiplicative mapping relationship between user attributes and language indicators.

[0124] In video interaction mode, image processing tools such as OpenCV can be used to detect the status of checkboxes in screen sharing in real time, and handwriting recognition algorithms can be used to identify annotation content. In conjunction with natural language processing models, cosine similarity matching is performed between annotations and the current page's text to generate an annotation relevance score. During data fusion, image frames and audio content are linked using a timestamp synchronization mechanism to ensure semantic context consistency.

[0125] In text interaction mode, a multi-classification model is used to label message types, combining rule templates with semantic vector distribution for dual-channel judgment. Urgency keyword matching utilizes a dual recognition strategy based on a lexicon and sentiment annotation. Sentence complexity is calculated through syntax tree depth or dependency graph path analysis. Attribute weighting factors are selected based on a configurable rule base, such as multiplying 1.2 for ages 30-50 and 1.3 for occupations such as teachers. The system automatically integrates these factors and outputs an intent score.

[0126] Example: In the healthcare sector, a user initiates a voice inquiry regarding chronic disease management services. The user's voice content is transcribed as "I have high blood pressure. Does this package include blood pressure medication?" The system identifies the interrogative sentence structure and the keyword "blood pressure medication." The interrogative word density reaches the high threshold, and the tone intensity, combined with intonation analysis, is marked as medium. The user's occupation field is "Worker" and their age is 52. The system combines the configured weighting factors and outputs a current user intent score of 2.1, marking it as high priority.

[0127] In a financial services scenario, a user checked the "Critical Illness Insurance Disclaimer" in an insurance demonstration video and hand-wrote "Does it cover stroke?" next to it. The system identified a high degree of semantic overlap between this annotation and the clause, with a correlation of 0.88 and a frequency of 1, resulting in a calculated video intent score of 1.3. Taking into account the user's family of two children and occupation as a physician, the system applied a factor multiplication and output an intent score of 1.9, which was used to activate the critical illness product description module.

[0128] This embodiment, by constructing a diversified intent analysis mechanism based on interaction types, can fully adapt to the expression characteristics of different information modalities, thereby having stable intent recognition capabilities when faced with input content such as spoken, graphical, and structured text. The introduction of current user information enhances the personalization and scenario adaptability of the scoring results, which not only improves the accuracy of semantic reasoning, but also provides quantifiable and comparable input indicators for subsequent demand generation and strategy matching. The overall design ensures that in the context of different information sources and diverse expression modes, it can stably output structured intent scores, driving efficient closed-loop downstream logic processing.

[0129] In one embodiment, the above step S40 includes:

[0130] S401, if the interaction type is a voice conversation, the requirement category template is set to a consultation requirement; if the interaction type is a video conversation, the requirement category template is set to a confirmation requirement; if the interaction type is a text conversation, the requirement category template is set to an operation requirement;

[0131] S402, extracting the occupation field in the current user information, and matching predefined industry-specific requirements according to the occupation field;

[0132] S403, extracting the family structure field in the current user information, and matching predefined associated requirement items according to the family structure field;

[0133] S404, dividing the current user intention score into a high score interval, a medium score interval, and a low score interval;

[0134] S405: Mark the higher score interval as an urgent priority, mark the medium score interval as a high priority, and mark the lower score interval as a normal priority;

[0135] S406, structurally encapsulate the requirement category template, the industry-specific requirement items, the associated requirement items and priorities to generate a structured requirement list including requirement category template fields, industry-specific requirement item fields, associated requirement item fields and priority fields.

[0136] In this embodiment, in the process of generating a requirements list, it is first necessary to match the appropriate requirements category template according to the interaction type. The interaction type is identified based on the user's interactive behavior pattern, such as voice, video or text. The voice interaction type is usually related to the user's consulting needs, so its requirement category template is set to "consulting type requirements". The video interaction type often includes confirmation or modification of specific content, so its requirement category template is set to "confirmation type requirements". The text interaction type usually indicates that the user has clearly expressed the specific operation intention, so it is set to "operation type requirements". This classification helps to refine the user's needs into specific categories according to different interaction types, thereby providing clear guidance for subsequent processing and response.

[0137] In terms of extracting user information, the system will extract the user's occupation field from the text corpus and use this information to match predefined industry-specific demand items. For example, if the user's occupation is "teacher", the corresponding industry-specific demand item may be "education fund insurance". If the user's occupation is "doctor", "critical illness insurance" may be matched as an industry-specific demand item. This process relies on the identification of the occupation field and its predefined mapping relationship with the demand items. Similarly, the system will also extract the family structure field from the current user information and match the associated demand items based on this information. For example, if the user has multiple children, the system can automatically match them with the "child medical insurance" demand item. This will help to accurately meet the specific needs of different users through personalized demand configuration.

[0138] The system then prioritizes the user's intent score according to pre-set rules. For example, high intent scores (1.5 or higher) are marked as "urgent priority," medium scores (1.0 or lower, <1.5) are marked as "high priority," and low intent scores are marked as "regular priority." This prioritization helps the system allocate resources and prioritize responses appropriately during subsequent policy generation.

[0139] Finally, all of this information—requirement category templates, industry-specific requirements, related requirements, and priorities—is encapsulated into a structured data format, typically in JSON. This structured requirements list includes multiple fields: requirement category template field (e.g., "consulting requirements"), industry-specific requirement field (e.g., "education fund insurance requirements"), related requirement field (e.g., "children's medical insurance"), and priority field (e.g., "urgent"). This structured data output enables the system to effectively provide precise input for subsequent operations, strategy generation, and other steps.

[0140] In scenarios that include all three types of interactions: voice, video, and text, the demand generation process needs to be processed separately based on the characteristics of each data source and then integrated to ensure the integrity of the semantic analysis and the accuracy of the generated results. The system first identifies consulting needs expressed by the user through transcription and sentence analysis based on the voice interaction content. For example, questions raised by the user through voice, such as "I want to know which diseases are covered by critical illness insurance", are identified as consulting-oriented demand items through the model. Subsequently, the system processes the video interaction content and identifies confirmation needs through the text recognition results in the keyframe images and the user's actions such as checking and handwritten annotations on the shared screen. For example, the user's annotation "Need to confirm whether it is effective" in the disclaimer position will be marked as confirmation-type attention content. For the text interaction part, the system identifies operational needs through structured processing. For example, the message sent by the user "Please change the policy address to Beijing" can be directly parsed as an operation execution request.

[0141] After the above three processing steps are completed, the system will represent each type of demand in a structured form and then enter the fusion stage. The fusion process will take into account multiple factors, including the degree of semantic overlap between demand items, the relative weight of intent scores, the credibility level of the source channels, etc. The system will deduplicate, merge and sort demand items from different sources, and label their corresponding demand category attributes based on the original interaction type, while further optimizing the final output based on the current user information. For example, if the terms confirmed by the user in the video are strongly correlated with the question items expressed in the voice, the system will integrate them into the same demand topic through semantic merging to avoid redundancy and improve strategy accuracy.

[0142] The final output, a structured requirements list, will include integrated consultation, confirmation, and operational requirements. Each item will be labeled with its source interaction type, semantic label, and priority level. This enables unified generation of requirements from multimodal input and provides semantic support for the precise formulation of subsequent interaction strategies. This approach effectively adapts to the complexity of requirements generation in multi-source, multimodal interaction scenarios, achieving decoupling of data sources and integration of semantic results.

[0143] This embodiment combines interaction type, user information, and intent scores to more accurately generate personalized request lists for each user. The structured output of request lists provides more efficient automated processing, reduces manual intervention, and improves the speed and accuracy of customer service responses. By prioritizing, the system intelligently allocates resources and adjusts service priorities, ensuring that the most needed requests receive timely responses.

[0144] In one embodiment, the above step S50 includes:

[0145] S501, constructing a user behavior graph based on historical interaction records;

[0146] S502, extracting collaborative association weights of interactive service items matching the demand list in the user behavior graph;

[0147] S503, sorting the demand items in the demand list based on the collaborative association weights of the interactive service items to generate an optimized demand priority list;

[0148] S504: Generate an optimized interaction strategy output based on the optimized demand priority list and the current user information.

[0149] In this embodiment, when generating interaction strategy outputs based on current user information and a list of needs, the system first constructs a user behavior graph based on historical interaction records, mapping the relationship structure between all users and various interactive services. This graph consists of multiple user nodes, interactive service nodes, and the associated edges between them. User nodes identify individual users, and interactive service nodes represent the various services provided by the platform. The edges connecting user nodes and service nodes carry attribute information such as usage frequency and time decay, which is used to calculate the strength of user preferences for specific services.

[0150] After the graph is established, the system further extracts the collaborative association weights between the service nodes related to the current user's generated wish list. The collaborative association weight is calculated based on the frequency of multiple services being used together in the same time period or scenario among all users. It can reflect the co-occurrence relationship and potential matching patterns between different services in the user group. The calculation method of this weight may include normalization of co-occurrence counts, time window weighting, service category similarity fusion, and other methods to improve the expressiveness of the association calculation.

[0151] After extracting the synergistic weights, the system will use the current user's wish list as a basis and sort the items in the list according to their synergy strength with other high-frequency services, resulting in an optimized priority list. This ranking not only considers the current user's intention strength and behavior path, but also incorporates the service pairing patterns found in group behavior, thereby generating more accurate and dynamic priority recommendations.

[0152] Finally, the system combines personalized attributes in the current user information, such as age group, occupation category, geographical distribution, historical preferences, etc., to further adjust the sorting strategy for the optimized demand priority list. For example, the ranking weight of health management services for doctor users can be increased, and the ranking priority of children's education security services for users from large families can be enhanced. By combining behavioral graph analysis with user portrait features, the system can output a multi-dimensional, dynamic interaction strategy for the current user, clearly recommending the user to pay attention to or execute the service process first, and providing highly adaptable strategic support for subsequent content push, interaction guidance, resource allocation, etc. This strategy is not only oriented towards current demand expressions, but also has a certain degree of predictiveness and linkage, and is suitable for automatic service decision generation in multiple strong situational scenarios such as medical and health management, financial and insurance consulting, etc.

[0153] In specific implementation, the behavior graph can be completed through a graph database or vector storage system, where nodes can use hash codes to represent user identity and service item types, and the attribute values of the edges can be generated based on the number of service calls within the time window, the time interval since the last access, and the access frequency. A sliding time window is used to count the usage records between each user node and the service item node, and the time decay coefficient is calculated using an exponential decay function, so that recent access behaviors have a higher weight in the graph. The calculation of collaborative association weights can be based on the normalized value obtained by dividing the co-occurrence frequency of multiple services in the same user history by the total number of users, or by integrating the label hierarchical structure between service items to improve the accuracy of semantic matching.

[0154] When optimizing demand priorities, a two-factor sorting mechanism based on demand intent score and synergy weight can be set. The system prioritizes the demand items according to the current user intent score, and then re-sorts them based on the synergy association weight of each demand item with other service items in the graph. The sorting strategy here can adopt a weighted comprehensive model with adjustable weights, in which the weight ratio of user intent score and synergy weight can be dynamically adjusted according to the user's profile category. For example, users with high behavioral stability are weighted towards intent score, while users with changeable preferences can increase the weight of the synergy graph.

[0155] The final policy output can be a structured service recommendation list, interaction path planning suggestions, or a generated, embeddable and executable service workflow file. For user behavior processes with long service paths, a visual interaction path diagram can also be output to help back-end decision-making systems make multi-step contact arrangements.

[0156] Example: In a healthcare business scenario, a user completes a remote video conversation with an agent. The generated structured list of requirements includes "family health consultation," "physical examination appointment service," and "chronic disease management tracking." The user behavior map constructed by the system shows that most users who choose the "chronic disease management tracking" service also frequently co-occur with "nutrition plan customization" and "family doctor service," and these requests are highly consistent with the current user profile. Based on this, the system adjusts the priority of the requirements, prioritizing "family health consultation" and recommending that the interaction strategy guide users to the nutrition plan module. The output strategy path is: "physical examination appointment → family health consultation → nutrition plan recommendation → chronic disease management tracking."

[0157] In the financial insurance sector, a user expressed interest in children's education insurance products through voice chat, generating a list of needs: "education fund insurance," "long-term critical illness insurance," and "family asset allocation." The system detected that the user profile was a teacher from a large family, and that this matched the frequently occurring "education fund insurance → supplementary accident insurance" combination in the behavioral graph. Therefore, the system added a recommendation to insert supplementary insurance into the strategy generation and increased the service guidance weight of "education fund insurance," resulting in the final strategy: "education fund insurance (primary) → supplementary accident insurance → family asset allocation recommendations."

[0158] This embodiment links current user information and wish lists with the service association structure in the user behavior graph, effectively optimizing policy output by combining user preferences with group behavior characteristics. This effectively improves the response accuracy and service matching of recommendation paths, avoiding service deviations caused by ambiguous user expressions or a single data modality. It also possesses contextual adaptation capabilities, dynamically adjusting sorting logic based on user identity attributes, achieving a dual-driven approach based on user profiles and service graphs. Ultimately, this achieves the combined effect of improving the targeted nature of recommended content, reducing the frequency of ineffective engagement, and enhancing user conversion efficiency.

[0159] In one embodiment, the above step S501 includes:

[0160] S5011, extracting user identifiers and interactive service item identifiers from historical interaction records of multiple users;

[0161] S5012, counting the number of interactions between each user identifier and the corresponding interactive service item identifier and the last interaction timestamp;

[0162] S5013, determining a time decay coefficient based on a time difference between the current time and the last interaction timestamp and the number of interactions;

[0163] S5014, determining a usage association weight between the user and the interactive service item based on the number of interactions and the time decay coefficient;

[0164] S5015, counting the number of co-occurrences of different interactive service item identifiers in the same user interaction record;

[0165] S5016, determining a collaborative association weight between interactive service items based on the co-occurrence count and the total number of users;

[0166] S5017: Generate a user behavior graph based on the usage association weight and the collaborative association weight.

[0167] In this embodiment, when constructing a user behavior graph, we first need to extract user identifiers and interactive service item identifiers from historical interaction data. User identifiers can be encrypted unique IDs, and service item identifiers can correspond to products, service processes, interactive interfaces, and other types. In actual systems, this data is often distributed in operation logs, interface call records, or customer service forms. Extraction and normalization logic is required to uniformly map heterogeneous sources into structured relationship pairs.

[0168] To establish preference relationships between users and services, it's necessary to extract user identifiers and service identifiers from multiple users' historical interaction records and perform multi-level statistical analysis based on the time dimension. First, using an event window mechanism, each user's interactions with each service are accumulated, recording the total number of interactions with the service and the time of the most recent interaction. Based on the interval between the current time and the most recent interaction, a time decay factor is calculated to dynamically adjust the influence of older behavior records. This decay factor can be calculated using a function model based on exponential decay, where the time difference is denoted as Δt and the decay sensitivity is determined by a preset parameter λ. It is expressed as the output of a natural exponential function with the product of negative λ and Δt as the exponent. This decay factor can be combined with the accumulated number of interactions to form a usage association weight that reflects the user's current level of activity with respect to a particular service. This weight dynamically measures the changing attractiveness of a service to a user, supporting personalized ranking in subsequent policy generation.

[0169] In addition, in order to characterize the common group connections between service items, it is necessary to count the co-occurrence behaviors of service pairs in the same user trajectory at the user granularity. The judgment of the co-occurrence relationship can be based on a set time overlap window. If the interaction time interval between two service items is less than a set threshold (for example, seven days), it is determined that they have co-occurred in the user behavior. By summarizing this statistical logic across multiple user dimensions, the number of co-occurrences of each pair of service items in the user group can be calculated. In order to eliminate the statistical bias caused by the difference in the number of users, the number of co-occurrences can be further divided by the total number of users in the system to generate the collaborative association weight between service items, reflecting their joint intention trends in the user group.

[0170] Based on the aforementioned usage association weights and collaboration association weights, a graph structure can be constructed to represent the multidimensional interaction network between users and services. The nodes in this graph include user identification nodes and service identification nodes, and the edge types are divided into usage association edges between user nodes and service nodes, and collaboration association edges between service nodes. This graph structure supports various operations such as path traversal, community detection, and edge weight sorting. It can be deployed in a graph database for persistent storage and can be incrementally and dynamically updated with real-time interaction records. It serves computational tasks such as matching priority scheduling, user path reasoning, and combination recommendation in the strategy generation module.

[0171] Example description: In the field of healthcare, the system monitors historical interactions between users and various health management service items, such as filling out health questionnaires, making appointments for chronic disease follow-up, and viewing physical examination reports. Suppose a user has opened the "Hypertension Health Follow-up" module three times in the past month, with the most recent interaction being 7 days ago. Assuming the time decay coefficient λ is 0.1 and the time difference Δt between the current time and the last interaction is 7, the time decay factor is exp(-0.1×7)≈0.4966. Combined with the cumulative number of interactions of 3, the calculated usage association weight is approximately 3×0.4966≈1.4898. This value indicates that the user still maintains a certain level of recent activity in the follow-up service and can be used as a priority candidate for recommending "continue follow-up" or "push health reminders" in subsequent strategy generation.

[0172] In the financial sector, for example, a user visited the "Critical Illness Insurance Plan" page five times in the past year, with the most recent visit occurring 15 days ago. With a total of 5 interactions, and λ set to 0.08, the decay factor is exp(-0.08 × 15) ≈ 0.3012, and the associated weight is 5 × 0.3012 ≈ 1.506. Although this user has a high cumulative number of interactions, their recent interest has decreased over time. Therefore, the system can appropriately lower their priority when generating policies unless they show interest again.

[0173] This embodiment constructs a graph structure by integrating interaction frequency, time decay, and service item co-occurrence relationships. This allows for the inclusion of group behavior patterns beyond individual user information, addressing information gaps caused by incomplete user expression or unclear identification of shifting interests. In actual policy recommendations, this graph not only enhances the adaptability of service ranking but also provides a natural transition structure between candidate services, providing a structured foundation for policy path planning.

[0174] In one embodiment, after the above step S504, the method further includes:

[0175] S505, extracting a sub-graph corresponding to the current user from the user behavior graph;

[0176] S506, analyzing the interaction window of the current user through a time series model based on the historical behavior path in the sub-graph;

[0177] S507: Update the usage association weight and the collaboration association weight in the user behavior graph according to the feedback behavior of the current user within the interaction window period.

[0178] In this embodiment, in order to continuously improve the dynamic adaptability of the output of the interaction strategy, a user behavior graph update and feedback closed-loop mechanism is further introduced on the basis of the existing optimized demand priority list. First, it is necessary to extract a subset of data related to the current user in the user behavior graph. The sub-graph structure includes the current user node and the interactive service item nodes associated with it in history. The edge weight information connecting these nodes is composed of usage association weights and collaborative association weights. Usage association weights are usually used to represent the degree of individual behavior closeness between a user and a service, while collaborative association weights reflect the potential combination preferences of multiple service items from the user's perspective. Through the sub-graph extraction process, it is possible to focus on the most relevant behavior paths and node relationships of the current user without loading the complete graph.

[0179] Subsequently, by analyzing the time series of historical behavior paths within the subgraph, a predictive model is used to determine the current user's interaction window in the future. This predictive model can be based on behavioral activity calculations within a sliding time window, interaction sequence modeling using a recurrent neural network, or a combination of short-term frequency and long-term silence patterns to determine the potential range of the window. Analysis of the interaction window helps identify whether the user's current intent is at a high-responsiveness stage, thereby driving real-time and pacing strategy generation.

[0180] When a user generates new interactive feedback within the window period, the user behavior graph should be partially updated. The update logic of the association weight is to combine the count value of the newly added interactive event with the timestamp of the event, and to make a weighted correction to the original weight according to the preset decay model, so that the activity intensity between the user and the service item remains timely. The update of the collaborative association weight is based on the combination of service items in the interactive feedback, re-counting the co-occurrence relationship of the service item pairs, and fine-tuning based on the current user's behavior to achieve two-way evolution of the graph between the group and individual levels.

[0181] Finally, the updated user behavior graph is used to regenerate the interaction strategy output. This output not only inherits the optimized demand priority list but also integrates recent user behavior trends with structured graph relationships, making more targeted adjustments in strategy ranking, content recommendation, response path planning, and other aspects, giving the system feedback perception and learning capabilities.

[0182] In practical applications, user behavior graphs can be constructed based on graph databases. Subgraph extraction can be performed by partitioning the graph using user node indexes and loading a subset of all edges and nodes associated with the user. Time series models can be selected using either an LSTM-based behavior prediction network or a fixed-length sliding window mechanism to analyze trends in user service access frequency over consecutive time periods to determine active periods. Behavioral feedback can be automatically collected from user operation logs. For example, behavioral events such as clicks, fills, stays, and forwards can all be converted into interaction signals. The update operation using association weights takes the following form: Assume the current weight is W0, and a new interaction occurs at time T1. Based on the current time T2 and the decay coefficient λ, the updated weight is W = W0 + exp(-λ × (T2 - T1)). Collaborative association weights can be updated using a local statistical coverage approach. This approach determines whether a user's single operation involves multiple services. If a predetermined time or task binding relationship is met, the co-occurrence count for this service pair is incremented by one and renormalized based on the total number of users.

[0183] The strategy output can be generated based on a ranking model or a rule engine. For example, in the ranking logic, the updated usage weight is used as a personalized scoring factor and combined with the original priority score to construct a weighted ranking expression, thereby improving the accuracy of personalized recommendations.

[0184] Example: In an insurance company's business process, agents collect interaction data from customer video conversations, phone recordings, and WeChat text messages. The system first identifies three types of interactions using metadata fields in the interaction data: video, voice, and text, and invokes corresponding parsing models for each. The voice data undergoes noise reduction and audio transcription to generate the first speech transcript. Keyframe images are extracted from the video data for OCR recognition and combined with the audio to generate the video parsed text. WeChat text data undergoes encoding and parsing, segmentation, and structured processing to generate structured text. All text forms undergo semantic alignment, and a standardized text corpus is constructed based on timestamps. Next, the system extracts user information from the standardized text corpus. For example, if the system identifies the sentence "I have two children and recently bought an apartment," it extracts "family structure = two children" and "regional preference = homebuyer." Combining the sentence "Can you recommend a product with more comprehensive coverage?" with the sentence "Please help me get it done as soon as possible," the system identifies it as an imperative sentence with a high urgency tone, resulting in a high speech intent score. Furthermore, based on the user's age (35), occupation (teacher), and family structure, a weighting factor is applied to determine the current user's intent score is high. Based on user information and intent scores, the system maps the interaction type to "video → confirmation request," occupation to "education fund insurance request," and family structure to "children's medical insurance request." Intent scores are assigned to urgency priorities, ultimately generating a structured list of requests. The system then retrieves the historical interaction graph and discovers that the user frequently inquires about education products and is interested in children's health insurance, matching these with the list of requests. The system calculates the co-occurrence frequency of these requests, calculates the synergistic association weights, and sorts them by weight, finding that education fund insurance has the highest priority. Based on the user's behavior of "considering comprehensive protection for their children," the system generates a recommended combination of "children's critical illness insurance + education fund + additional hospitalization medical care" as the optimized policy output. After the policy is pushed, the user clicks on the education fund insurance product and initiates the application process within the interaction window. The system detects this behavior and updates the usage weight of the education fund insurance. Furthermore, since the user also viewed the children's critical illness insurance, the synergistic association weight of the two is increased. The updated user behavior graph is stored in the graph database to support the next round of recommendations.

[0185] In hospital scenarios, patients interact with the platform through smart health terminals in a multimodal manner. Data sources for these interactions include voice consultations, medical image annotations, and online text-based consultations. After the system identifies the interaction type, the voice data is transcribed into a first-line speech transcript using a model. Medical image annotations and text descriptions are combined into video parsed text. The medical records from the text-based consultations are parsed into structured text. These three types of text corpora are then integrated and synchronized to generate standardized medical semantic representations. The system extracts patient information from the text corpus, such as "I have hypertension and just had a brain CT scan last year. My father has a similar medical history," and extracts fields such as "Chronic disease = hypertension," "Existing imaging records," and "Genetic risk = positive." After analyzing the anxious tone in the voice and the text "Please confirm a diagnosis as soon as possible," a high-intensity medical intent score is generated. Combining the patient's age (58 years old) and family medical history, a weighting factor for user attributes is determined to form the current user intent score. The system then generates a needs list, mapping the video interaction to a confirmation need and the medical history information to medical needs such as "hypertension complication screening" and "stroke warning," which are mapped to a high-priority intent score. The system constructs a medical knowledge graph and, based on the patient's historical medical records, analyzes the co-occurrence relationship between hypertension screening and neurological examinations, prioritizing the recommendation of the "Integrated Cardiovascular and Cerebrovascular Checkup" plan. If the user clicks on this checkup package and uploads a previous CT scan during the interaction window, the system updates the usage weight of this service item in the behavior graph based on this behavior. It also identifies the co-occurrence between the CT scan upload behavior and the "Brain Checkup" service, increasing their synergistic weight and simultaneously updating the graph node relationships to support the next precise recommendation process.

[0186] This embodiment combines sub-graph extraction, behavior window prediction, and a dynamic graph update mechanism based on demand priorities, enabling the system to continuously adjust service response strategies based on evolving user behavior. This not only improves the timeliness and relevance of responses, but also effectively enhances the system's adaptability to changes in individual behavior patterns. In particular, in scenarios with volatile user needs and complex service relationships, a flexible strategy reconstruction mechanism driven by a structured graph can be implemented, significantly improving the timeliness of traditional static recommendations.

[0187] In one embodiment, an intention analysis and policy generation device is provided, which corresponds one-to-one to the intention analysis and policy generation method in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of the device for intent analysis and strategy generation of the present invention. It includes an interaction recognition module 10, a multimodal analysis module 20, an intent scoring module 30, a demand generation module 40, and a strategy output module 5. Each functional module is described in detail below:

[0188] An interaction identification module 10 is used to obtain interaction data and identify the interaction type based on metadata features of the interaction data;

[0189] A multimodal parsing module 20 is configured to select a corresponding parsing model based on the interaction type, and parse the interaction data according to the parsing model to generate a text corpus;

[0190] An intention scoring module 30 is configured to extract current user information from the text corpus and perform user intention scoring analysis based on the intention analysis strategy corresponding to the interaction type to generate a current user intention score;

[0191] A demand generation module 40 is configured to generate a demand list based on the current user information, the current user intention score, and the interaction type;

[0192] The strategy output module 50 is configured to generate an interaction strategy output based on the current user information and the requirement list.

[0193] In one embodiment, the multimodal analysis module 20 is specifically configured to:

[0194] When the interaction type is a voice conversation, performing noise reduction processing on the audio data in the interaction data, extracting an audio track feature vector from the noise-reduced audio data, and performing phoneme alignment and context error correction on the audio track feature vector using a pre-trained speech recognition model to generate a first speech transcription text;

[0195] When the interaction type is a video conversation, extracting key frame images from the video data in the interaction data, identifying text content in the key frame images using a text recognition model, performing voice transcription on the audio track of the video data to generate a second voice transcription text, and aligning and fusing the text content with the second voice transcription text along a time axis to generate a video parsed text;

[0196] When the interaction type is a text conversation, identifying the encoding format of the original text in the interaction data, decoding and converting the original text based on the encoding format to generate text data in a unified encoding format, and segmenting the text data in the unified encoding format to generate structured conversation text;

[0197] Cross-modal semantic alignment is performed on the first speech transcription text, the video parsed text and / or the structured conversation text to generate a standardized text corpus.

[0198] In one embodiment, the intention scoring module 30 is specifically configured to:

[0199] When the interaction type is a voice conversation, extracting interrogative sentences and imperative sentences from the text corpus as demand description fields;

[0200] Analyze the question word density and tone intensity value in the requirement description field to generate a speech intent score;

[0201] Determining a user attribute weighting factor based on the occupation, age, and / or family structure in the current user information, and multiplying the voice intent score by the user attribute weighting factor to generate a current user intent score;

[0202] When the interaction type is a video conversation, identifying a check box and handwritten annotations in a screen sharing area in the text corpus as a requirement confirmation field;

[0203] Counting the frequency of annotation selections and the relevance of annotations in the requirement confirmation field to generate a video intent score;

[0204] Determining a user attribute weighting factor based on occupation, age and / or family structure in the current user information;

[0205] Multiplying the video intent score by the user attribute weighting factor to generate a current user intent score;

[0206] When the interaction type is a text conversation, classifying the text corpus into message types, and marking the demand urgency field in the text corpus according to the message type;

[0207] Counting the keyword matching degree and sentence complexity in the demand urgency field to generate a text intent score;

[0208] A user attribute weighting factor is determined based on the occupation, age and / or family structure in the current user information, and the text intent score is multiplied by the user attribute weighting factor to generate a current user intent score.

[0209] In one embodiment, the demand generation module 40 is specifically configured to:

[0210] If the interaction type is a voice conversation, the requirement category template is set to a consultation requirement; if the interaction type is a video conversation, the requirement category template is set to a confirmation requirement; if the interaction type is a text conversation, the requirement category template is set to an operation requirement;

[0211] Extracting the occupation field in the current user information, and matching predefined industry-specific requirement items according to the occupation field;

[0212] Extracting a family structure field from the current user information, and matching predefined associated requirement items according to the family structure field;

[0213] Dividing the current user intention score into a high score interval, a medium score interval, and a low score interval;

[0214] Mark the higher score interval as urgent priority, mark the medium score interval as high priority, and mark the lower score interval as regular priority;

[0215] The requirement category template, the industry-specific requirement items, the associated requirement items and priorities are structurally packaged to generate a structured requirement list including requirement category template fields, industry-specific requirement item fields, associated requirement item fields and priority fields.

[0216] In one embodiment, the policy output module 50 is specifically configured to:

[0217] Build user behavior graphs based on historical interaction records;

[0218] Extracting collaborative association weights of interactive service items matching the demand list from the user behavior graph;

[0219] Sorting the demand items in the demand list based on the collaborative association weights of the interactive service items to generate an optimized demand priority list;

[0220] An optimized interaction strategy output is generated based on the optimized demand priority list and the current user information.

[0221] In one embodiment, the policy output module 50 is specifically configured to:

[0222] Extracting user identifiers and interactive service item identifiers from historical interaction records of multiple users;

[0223] Count the number of interactions between each user identifier and the corresponding interactive service item identifier and the timestamp of the last interaction;

[0224] determining a time decay coefficient based on a time difference between a current time and a last interaction timestamp and the number of interactions;

[0225] Determining a usage association weight between the user and the interactive service item based on the number of interactions and the time decay coefficient;

[0226] Count the number of co-occurrences of different interactive service item identifiers in the same user interaction record;

[0227] Determining collaborative association weights between interactive service items based on the co-occurrence count and the total number of users;

[0228] A user behavior graph is generated according to the usage association weight and the collaborative association weight.

[0229] In one embodiment, the policy output module 50 is specifically configured to:

[0230] Extracting a sub-graph corresponding to the current user from the user behavior graph;

[0231] Based on the historical behavior path in the sub-graph, the interaction window of the current user is analyzed through a time series model;

[0232] According to the feedback behavior of the current user within the interaction window period, the usage association weight and the collaboration association weight in the user behavior map are updated.

[0233] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the service side of an intent analysis and policy generation method.

[0234] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the user side of an intention analysis and strategy generation method.

[0235] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0236] Acquire interaction data, and identify interaction types based on metadata features of the interaction data;

[0237] Selecting a corresponding parsing model based on the interaction type, and parsing the interaction data according to the parsing model to generate a text corpus;

[0238] Extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score;

[0239] generating a requirements list according to the current user information, the current user intention score, and the interaction type;

[0240] Generate an interaction strategy output based on the current user information and the requirement list.

[0241] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0242] Acquire interaction data, and identify interaction types based on metadata features of the interaction data;

[0243] Selecting a corresponding parsing model based on the interaction type, and parsing the interaction data according to the parsing model to generate a text corpus;

[0244] Extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score;

[0245] generating a requirements list according to the current user information, the current user intention score, and the interaction type;

[0246] Generate an interaction strategy output based on the current user information and the requirement list.

[0247] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0248] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0249] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0250] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for intention analysis and strategy generation, characterized in that: The following steps are involved: Acquire interaction data, and identify interaction types based on metadata features of the interaction data; Selecting a corresponding parsing model based on the interaction type, and parsing the interaction data according to the parsing model to generate a text corpus; Extracting current user information from the text corpus, and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score; generating a requirements list according to the current user information, the current user intention score, and the interaction type; Generate an interaction strategy output based on the current user information and the requirement list.

2. The intention analysis and strategy generation method according to claim 1, wherein: Selecting a corresponding parsing model based on the interaction type, and parsing the interaction data according to the parsing model to generate text corpus, including: When the interaction type is a voice conversation, performing noise reduction processing on the audio data in the interaction data, extracting an audio track feature vector from the noise-reduced audio data, and performing phoneme alignment and context error correction on the audio track feature vector using a pre-trained speech recognition model to generate a first speech transcription text; When the interaction type is a video conversation, extracting key frame images from the video data in the interaction data, identifying text content in the key frame images using a text recognition model, performing voice transcription on the audio track of the video data to generate a second voice transcription text, and aligning and fusing the text content with the second voice transcription text along a time axis to generate a video parsed text; When the interaction type is a text conversation, identifying the encoding format of the original text in the interaction data, decoding and converting the original text based on the encoding format to generate text data in a unified encoding format, and segmenting the text data in the unified encoding format to generate structured conversation text; Cross-modal semantic alignment is performed on the first speech transcription text, the video parsed text and / or the structured conversation text to generate a standardized text corpus.

3. The intention analysis and strategy generation method according to claim 1, wherein: Extracting current user information from the text corpus and performing user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score includes: When the interaction type is a voice conversation, extracting interrogative sentences and imperative sentences from the text corpus as demand description fields; Analyze the question word density and tone intensity value in the requirement description field to generate a speech intent score; Determining a user attribute weighting factor based on the occupation, age, and / or family structure in the current user information, and multiplying the voice intent score by the user attribute weighting factor to generate a current user intent score; When the interaction type is a video conversation, identifying a check box and handwritten annotations in a screen sharing area in the text corpus as a requirement confirmation field; Counting the frequency of annotation selections and the relevance of annotations in the requirement confirmation field to generate a video intent score; Determining a user attribute weighting factor based on occupation, age and / or family structure in the current user information; Multiplying the video intent score by the user attribute weighting factor to generate a current user intent score; When the interaction type is a text conversation, classifying the text corpus into message types, and marking the demand urgency field in the text corpus according to the message type; Counting the keyword matching degree and sentence complexity in the demand urgency field to generate a text intent score; A user attribute weighting factor is determined based on the occupation, age and / or family structure in the current user information, and the text intent score is multiplied by the user attribute weighting factor to generate a current user intent score.

4. The intention analysis and strategy generation method according to claim 1, wherein: Generating a requirements list according to the current user information, the current user intention score, and the interaction type, including: If the interaction type is a voice conversation, the requirement category template is set to a consultation requirement; if the interaction type is a video conversation, the requirement category template is set to a confirmation requirement; if the interaction type is a text conversation, the requirement category template is set to an operation requirement; Extracting the occupation field in the current user information, and matching predefined industry-specific requirement items according to the occupation field; Extracting a family structure field from the current user information, and matching predefined associated requirement items according to the family structure field; Dividing the current user intention score into a high score interval, a medium score interval, and a low score interval; Mark the higher score interval as urgent priority, mark the medium score interval as high priority, and mark the lower score interval as regular priority; The requirement category template, the industry-specific requirement items, the associated requirement items and priorities are structurally packaged to generate a structured requirement list including requirement category template fields, industry-specific requirement item fields, associated requirement item fields and priority fields.

5. The intention analysis and strategy generation method according to claim 1, wherein: Generating an interaction strategy output based on the current user information and the requirement list includes: Build user behavior graphs based on historical interaction records; Extracting collaborative association weights of interactive service items matching the demand list in the user behavior graph; Sorting the demand items in the demand list based on the collaborative association weights of the interactive service items to generate an optimized demand priority list; An optimized interaction strategy output is generated based on the optimized demand priority list and the current user information.

6. The intention analysis and strategy generation method according to claim 5, characterized in that: Build a user behavior graph based on historical interaction records, including: Extracting user identifiers and interactive service item identifiers from historical interaction records of multiple users; Count the number of interactions between each user identifier and the corresponding interactive service item identifier and the timestamp of the last interaction; determining a time decay coefficient based on a time difference between a current time and a last interaction timestamp and the number of interactions; Determining a usage association weight between the user and the interactive service item based on the number of interactions and the time decay coefficient; Count the number of co-occurrences of different interactive service item identifiers in the same user interaction record; Determining collaborative association weights between interactive service items based on the co-occurrence count and the total number of users; A user behavior graph is generated according to the usage association weight and the collaborative association weight.

7. The intention analysis and strategy generation method according to claim 5, wherein: After generating an optimized interaction strategy output based on the optimized demand priority list and the user information, the method further includes: Extracting a sub-graph corresponding to the current user from the user behavior graph; Based on the historical behavior path in the sub-graph, the interaction window of the current user is analyzed through a time series model; According to the feedback behavior of the current user within the interaction window period, the usage association weight and the collaboration association weight in the user behavior map are updated.

8. An intention analysis and strategy generation device, characterized in that: The intention analysis and strategy generation device includes: An interaction identification module, configured to obtain interaction data and identify interaction types based on metadata features of the interaction data; A multimodal parsing module, configured to select a corresponding parsing model based on the interaction type, and parse the interaction data according to the parsing model to generate a text corpus; An intent scoring module is used to extract current user information from the text corpus and perform user intent scoring analysis based on the intent analysis strategy corresponding to the interaction type to generate a current user intent score; A demand generation module, configured to generate a demand list based on the current user information, the current user intention score, and the interaction type; A strategy output module is used to generate an interaction strategy output based on the current user information and the requirement list.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and an intent analysis and policy generation program stored in the memory and capable of running on the processor. When the intent analysis and policy generation program is executed by the processor, the steps of the intent analysis and policy generation method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores an intent analysis and policy generation program, which, when executed by a processor, implements the steps of the intent analysis and policy generation method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent chart dynamic generation method based on voice recognition and multi-modal interaction

    CN120670583A

  • Insurance product multi-dimensional matching method, system and equipment based on behavior portrait

    CN121032683A

  • Data use control strategy determination method and device, medium, equipment and product

    CN121144566A

  • Data analysis method and device, electronic equipment, readable storage medium and chip

    CN121145886A

  • Data analysis method and device, electronic equipment, readable storage medium and chip

    CN121145886B