Medical sales demand response method and system based on multi-mode AI intelligent system
By using the three-dimensional quality scoring model and spatiotemporal attention fusion model of the multimodal AI intelligent system, the problems of multimodal data fusion anti-interference and intent recognition in the pharmaceutical sales demand response system were solved, realizing the generation of personalized response content and dynamic optimization of the system, and improving the accuracy and adaptability of the response system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-27
AI Technical Summary
Existing pharmaceutical sales demand response systems suffer from weak anti-interference capabilities in multimodal data fusion, susceptibility to data quality fluctuations in intent recognition, and low adaptability of personalized recommendations to real-time sales scenarios, making it difficult to meet the efficient service needs of the pharmaceutical sales industry.
A multimodal AI intelligent system-based approach is adopted, which evaluates data quality through a three-dimensional quality scoring model, combines differentiated enhancement strategies and a spatiotemporal attention fusion model to generate personalized multimodal response content, and optimizes the core model of the system through user feedback data to achieve dynamic adaptation.
It significantly improves the accuracy and anti-interference capability of multimodal fusion, optimizes personalized response and scene adaptability, has continuous adaptation capability, and improves the stability of the response system and user experience.
Smart Images

Figure CN121745758A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of AI intelligence and medical sales cross, and particularly relates to a medical sales demand response method and system based on a multi-modal AI intelligent system. BACKGROUND
[0002] In the process of digital transformation of the medical sales industry, the presentation form of user demand is increasingly diversified, which has expanded from traditional single text consultation to multi-modal demand data such as text modalities, voice modalities, image modalities and sales scene environment signal modalities. Such multi-modal demand data has become the core input basis for realizing accurate demand response of medical sales.
[0003] However, the current existing medical sales demand response system still has many technical defects in adapting to the processing scene of multi-modal demand data, and it is difficult to meet the actual needs of efficient service of the industry: First, the anti-interference ability of multi-modal data fusion is weak. The existing system adopts simple splicing, weighted summation and other extensive fusion means for multi-modal data processing, without fully considering the semantic correlation degree between different modal data, and without differentiated processing according to the quality difference of each modal data. When there is noise, missing or interference in a modal data, the overall fusion feature is distorted, which directly affects the accuracy of subsequent demand identification.
[0004] Second, the intention recognition result is easily affected by data quality fluctuations. The multi-modal demand data collected in the medical sales scene has significant differences in completeness, clarity and timeliness, and the data quality is uneven. However, the existing system lacks a quality evaluation system and differentiated enhancement processing mechanism specifically for multi-modal data. The original data is directly input into the intention recognition model, which leads to poor stability of the judgment results of demand intention categories, priorities and urgency, and is difficult to adapt to the decision-making needs of complex medical sales scenes.
[0005] Third, the adaptation degree of personalized recommendation to real-time sales scene is low. The response content generation of the existing system depends on fixed templates or general recommendation algorithms, without deeply integrating real-time inventory data, regional medical policy differences and user personalized portraits and other key information. Moreover, the response modalities are mostly single text form, which cannot match the diversified interaction preferences of users, resulting in insufficient practicality of the recommended content and poor user experience.
[0006] Fourthly, the system lacks a dynamic optimization mechanism. The core model (such as the intention recognition model, the recommendation model, etc.) of the existing system is fixed after training, and a dynamic optimization closed loop based on user feedback data and real-time sales data (such as recommendation conversion rate, inventory turnover efficiency, etc.) is not established, which is difficult to adapt to the rapid changes of the medical sales market policy, inventory status and user demand, and the response accuracy of the system may decrease over time.
[0007] Therefore, it is a technical problem to be solved in the process of digital transformation of the current medical sales industry to develop a medical sales demand response scheme that can realize high-quality fusion of multi-modal data, accurate intention recognition, scenario-based personalized response and dynamic optimization, so as to improve the anti-interference ability, recognition accuracy and scenario adaptability of the response system. SUMMARY
[0008] The purpose of the present application is to provide a medical sales demand response method and system based on a multi-modal AI intelligent system, which solves the technical problems of weak anti-interference ability of multi-modal fusion, intention recognition easily affected by data quality fluctuations, and low adaptability of personalized recommendation and real-time sales scenario in the prior art.
[0009] To achieve the above purpose, the present application provides a medical sales demand response method based on a multi-modal AI intelligent system, comprising the following steps: S1: acquiring multi-modal demand data, inputting the demand data of each modality into a pre-constructed three-dimensional quality scoring model, and outputting the data quality score of the corresponding modality from the three-dimensional quality scoring model; wherein the multi-modal at least includes: text modality, voice modality, image modality and sales scene environment signal modality; S2: based on the data quality score of each modality, a differentiated enhancement strategy is adopted for the demand data of the corresponding modality, and an enhanced multi-modal feature is output; S3: input the enhanced multi-modal feature and the data quality score of the corresponding modality into the pre-set spatio-temporal attention fusion model, and output the multi-modal fusion feature from the spatio-temporal attention fusion model; then input the multi-modal fusion feature into the pre-set deep learning model, and output the demand intention category, priority and urgency score from the deep learning model; S4: combining real-time inventory data, regional medical policy and user portrait information, generating personalized multi-modal response content adapted to the scene based on the demand intention category, priority and urgency score; S5: based on the collected user feedback data on the personalized multi-modal response content adapted to the scene and real-time sales data, optimizing the system core model to obtain an optimized system core model; the system core model at least includes: spatio-temporal attention fusion model and deep learning model.
[0010] As described above, the sub-steps for pre-constructing the three-dimensional quality scoring model are as follows: S11: Combining the application requirements of multimodal demand data in the pharmaceutical sales scenario, by analyzing the impact of historical multimodal demand data on the two core links of demand identification and response generation, determine the initial evaluation dimensions. The initial evaluation dimensions should at least include data completeness, clarity, and timeliness; S12: Based on the initial evaluation dimensions, access the pre-constructed evaluation dimension importance judgment database and select the importance judgment data corresponding to the initial evaluation dimensions as target data; use the analytic hierarchy process (AHP) to determine the weight coefficients of each initial evaluation dimension based on the target data; S13: Based on the initial evaluation dimensions and their weight coefficients, construct the initial three-dimensional quality scoring model; S14: Based on multimodal sample data labeled with quality tags, optimize and validate the initial three-dimensional quality scoring model to obtain the final three-dimensional quality scoring model.
[0011] As shown above, the initial three-dimensional quality scoring model includes a quantization processing layer and a quality scoring layer. The quantization processing layer receives the demand data for each modality and performs quantification processing on the completeness, clarity, and timeliness indicators of the demand data for each modality to obtain the quantified values of completeness, clarity, and timeliness of the demand data for each modality. These values are then input into the quality scoring layer, which outputs the data quality score corresponding to the demand data for each modality.
[0012] As shown above, the expression for the data quality score corresponding to the demand data of each modality output by the quality scoring layer is: ; in, For the first The data quality score corresponding to the demand data for each modality; For integrity weight, For clarity weight, As a weight for timeliness, ; For the first The completeness quantification value of the demand data for each modality; For the first Clear quantifications of demand data for each modality; For the first The timeliness quantification value of demand data for each modality.
[0013] The sub-step of outputting the enhanced multi-modal feature based on the data quality score of each modality, and adopting a differential enhancement strategy for the demand data of the corresponding modality is as follows: S21: according to the data quality score of each modality, traversing a preset quality level and enhancement strategy corresponding table to determine the quality level corresponding to the demand data of each modality and the matched enhancement strategy; the quality level and enhancement strategy corresponding table sets rules for multiple modalities respectively, each modality corresponds to three quality level range values, one quality level range value corresponds to one quality level, and one quality level corresponds to one enhancement strategy; S22: based on the matched enhancement strategy of the demand data of each modality, the demand data of the corresponding modality is enhanced to obtain the enhanced features of the demand data of each modality, and the enhanced features of the demand data of each modality at least include enhanced text features, enhanced voice features, enhanced image features and enhanced sales scene environment signal features; S23: the enhanced features of the demand data of each modality are cross-modal aligned and structuredly fused according to the semantic correlation degree of the medical sales scene to obtain structured multi-modal fusion features; S24: the structured multi-modal fusion features are checked for feature integrity and scene adaptability, wherein the feature integrity check is to check the coverage of the key information of each modality in the structured multi-modal fusion features, and if the coverage is greater than or equal to a coverage threshold, it is determined that the check is passed, and the scene adaptability check is to match the structured multi-modal fusion features with the typical demand features of the medical sales scene, and if the matching degree is greater than or equal to a matching degree threshold, it is determined that the check is passed; if both checks are passed, the structured multi-modal fusion features are output as the final enhanced multi-modal features, and if any check is not passed, S21 is re-executed.
[0014] As above, wherein the coverage threshold = 90%, and the matching degree threshold = 95%.
[0015] As above, wherein, in combination with real-time inventory data, regional medical policies and user portrait information, based on demand intent categories, priority and urgency scores, the sub-step of generating personalized multi-modal response content adapted to the scene is as follows: S41: structurally cleaning the extracted real-time inventory data, regional medical policies and user portrait information, and the integrated demand intent categories, priority and urgency scores to form a standardized parameter data set; S42: calculating the standardized parameter data set through a preset recommendation priority model to obtain the recommendation priority of each response content, and each response content at least includes drug recommendation, policy interpretation and drug purchase guide; S43: based on the recommendation priority and the standardized parameter data set, generating personalized multi-modal response content adapted to the scene according to the user interaction preference in the user portrait information.
[0016] The substep of generating the personalized multi-modal response content adapted to the scene based on the user interaction preference in the user portrait information and the recommendation priority and the standardized parameter dataset is as follows: S431: based on the recommendation priority, the response contents are sorted in descending order to generate a response content display sequence, and the response contents with high recommendation priority are preferentially displayed; S432: according to the priority sorting result of the response content display sequence, the first N response contents corresponding to the recommendation priority are selected as target response contents, and N is a preset positive integer; S433: an existing multi-modal generation tool is called to determine a target presentation mode of the target response content according to the user interaction preference in the user portrait information, wherein the user interaction preference at least includes one or more of a text presentation mode, a voice presentation mode and a graphic-text presentation mode; S434: according to the target response content, a corresponding response content template is called, and real-time inventory data, regional medical policy, user portrait information, demand intention category, priority and urgency score in the standardized parameter dataset are filled into the response content template according to the format requirement of the target presentation mode to obtain an initial multi-modal response content; S435: through an existing data verification tool, the filled information in the initial multi-modal response content is compared and verified with the standardized parameter dataset to ensure that the filled information is unbiased, and the initial multi-modal response content is output as the personalized multi-modal response content adapted to the scene.
[0017] The substep of optimizing the system core model based on the collected feedback data of the user on the personalized multi-modal response content adapted to the scene and the real-time data of the sales link is as follows: S51: the feedback data of the user on the personalized multi-modal response content adapted to the scene and the real-time data of the sales link are collected, and the feedback data at least includes a user satisfaction score and a correction suggestion, and the real-time data of the sales link at least includes a recommendation conversion rate and an inventory turnover efficiency; S52: the feedback data and the real-time data of the sales link are input into a preset comprehensive optimization model, and a comprehensive optimization coefficient is output by the comprehensive optimization model; S53: the comprehensive optimization coefficient is analyzed by using a preset two-stage optimization threshold, when the comprehensive optimization coefficient is greater than or equal to the two-stage optimization threshold, the current parameters and structure of the system core model are maintained unchanged, and the feedback data and the real-time data of the sales link are continuously collected to dynamically monitor the system performance; when the comprehensive optimization coefficient is less than the two-stage optimization threshold, two-stage optimization is triggered to obtain an optimized system core model, and the two-stage optimization includes parameter-level optimization and structure-level optimization; the parameter-level optimization includes updating the parameters in the spatio-temporal attention fusion model; the structure-level optimization includes adjusting the layer structure of the deep learning model through neural architecture search.
[0018] The application also provides a medicine sales demand response system based on a multi-modal AI intelligent system, comprising a data acquisition subsystem, a medicine sales demand response center and a user terminal; wherein the data acquisition subsystem is used for acquiring multi-modal demand data and transmitting the multi-modal demand data to the medicine sales demand response center; meanwhile, feedback data of user's personalized multi-modal response content to the adaptive scene and real-time data of the sales link are acquired and sent to the medicine sales demand response center; the medicine sales demand response center is used for executing the above-mentioned medicine sales demand response method based on the multi-modal AI intelligent system; and the user terminal is used for providing a multi-modal demand submission portal for the user and receiving and displaying the personalized multi-modal response content of the adaptive scene output by the medicine sales demand response center.
[0019] The application achieves the following beneficial effects: (1) The medicine sales demand response method and system based on the multi-modal AI intelligent system can significantly improve the multi-modal fusion precision and anti-interference capability: the quality of the demand data of each mode is quantitatively evaluated by a three-dimensional quality scoring model, the low-quality data is optimized by a differentiated enhancement strategy, the deep semantic fusion of multi-modal features is realized by a spatio-temporal attention fusion model, the interference of low-quality data is effectively avoided, and the stability and accuracy of intent recognition are greatly improved.
[0020] (2) The medicine sales demand response method and system based on the multi-modal AI intelligent system can comprehensively optimize personalized response and scene adaptability: core information such as real-time inventory data, regional medicine policies and user portrait information is deeply integrated, multi-modal response forms are determined based on recommendation priority calculation and user interaction preferences, so that the response content not only fits the real-time sales scene dynamics, but also accurately matches the personalized needs of users, significantly improving the practicality of the response content and user experience.
[0021] (3) The medicine sales demand response method and system based on the multi-modal AI intelligent system has a sustained adaptation capability: by collecting feedback data of user's personalized multi-modal response content to the adaptive scene and real-time data of the sales link, a two-stage optimization mechanism of the core model of the system is constructed, dynamic adjustment of the model parameters and structure is realized, and the response accuracy of the system can be continuously adapted to changes in market policies, inventory status and user needs in the long-term use process, thereby prolonging the effective service period of the system. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, brief descriptions will be given to the drawings needed to be used in the embodiments or prior art descriptions. Obviously, the drawings in the following description are only some embodiments described in the application, and other drawings can also be obtained by those skilled in the art based on these drawings.
[0023] Figure 1 Structure diagram of one embodiment of a medical sales demand response system based on a multi-modal AI intelligent system; Figure 2 Flowchart of one embodiment of a medical sales demand response method based on a multi-modal AI intelligent system. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0025] As shown in the drawings, Figure 1 The present application provides a medical sales demand response system based on a multi-modal AI intelligent system, comprising: a data acquisition subsystem 1, a medical sales demand response center 2, and a user terminal 3.
[0026] The data acquisition subsystem 1 is used to acquire multi-modal demand data and transmit the multi-modal demand data to the medical sales demand response center. At the same time, feedback data of user's personalized multi-modal response content to the adaptive scene and real-time data of the sales link are collected and sent to the medical sales demand response center.
[0027] The medical sales demand response center 2 is used to execute the following medical sales demand response method based on a multi-modal AI intelligent system.
[0028] The user terminal 3 is used to provide a multi-modal demand submission portal for users, and receive and display the personalized multi-modal response content of the adaptive scene output by the medical sales demand response center.
[0029] As shown in the drawings, Figure 2 The present application provides a medical sales demand response method based on a multi-modal AI intelligent system, comprising the following steps: S1: acquiring multi-modal demand data, inputting the demand data of each modality into a pre-constructed three-dimensional quality scoring model, and outputting the data quality score of the corresponding modality from the three-dimensional quality scoring model; wherein the multi-modal at least includes: text modality, voice modality, image modality and sales scene environment signal modality.
[0030] Further, the sub-steps of pre-constructing the three-dimensional quality scoring model are as follows: S11: In combination with the application requirements of multi-modal demand data in the medical sales scene, the influence degree of multi-modal historical demand data on the two core links of demand identification and response generation is analyzed to determine the initial evaluation dimensions, and the initial evaluation dimensions at least include data integrity, clarity and timeliness.
[0031] Specifically, all modal demand data are uniformly judged in quality by using the same set of evaluation dimensions, which can effectively avoid the quality score deviation caused by the difference of evaluation standards of different modes, and further improve the accuracy of data quality score and the comparability of quality of each modal data.
[0032] S12: Based on the initial evaluation dimensions, access the pre-constructed evaluation dimension importance judgment database, select the importance judgment data corresponding to the initial evaluation dimensions as the target data; use the analytic hierarchy process (AHP) to determine the weight coefficient of each initial evaluation dimension based on the target data.
[0033] Specifically, the pre-constructed evaluation dimension importance judgment database has the core data that a plurality of experts in the field of medical sales and AI data processing use the industry common 1-9 scale method (1 represents that two dimensions are equally important, 3 represents that one dimension is slightly more important than another dimension, 5 represents obvious importance, 7 represents strong importance, 9 represents extreme importance, and 2, 4, 6 and 8 are intermediate values of the adjacent judgments) to score the two-by-two comparison results of each initial evaluation dimension. When constructing the database, first collect the two-by-two comparison scoring data of the experts, and then arrange the scoring data into a judgment matrix corresponding to the initial evaluation dimensions; perform consistency check (consistency ratio CR<0.1) on the judgment matrix, and after verification, store the judgment matrix and the two-by-two comparison scoring data (i.e. original scoring data) corresponding thereto in the evaluation dimension importance judgment database to form standardized data resources that can be called for subsequent weight calculation. In addition, the analytic hierarchy process can be realized by using existing mature technology, so it is not described again; the weight coefficients of each initial evaluation dimension obtained by the analytic hierarchy process satisfy wherein, is the integrity weight; is the clarity weight; is the timeliness weight.
[0034] S13: Based on the initial evaluation dimensions and the weight coefficients of the initial evaluation dimensions, an initial three-dimensional quality score model is constructed.
[0035] Further, the initial three-dimensional quality score model comprises: a quantization processing layer and a quality score layer; the requirement data of each modality is received by the quantization processing layer, and the integrity index, the clarity index and the timeliness index of the requirement data of each modality are quantized respectively to obtain the integrity quantization value, the clarity quantization value and the timeliness quantization value of the requirement data of each modality, and are input to the quality score layer, and the data quality score corresponding to the requirement data of each modality is output by the quality score layer.
[0036] The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: ; The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: ; The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is: The expression of the data quality score corresponding to the requirement data of each modality output by the quality score layer is:
[0037] Further, the integrity index (such as the proportion of missing text characters, the proportion of valid speech duration), the clarity index (such as text recognition degree, image resolution, speech signal-to-noise ratio) and the timeliness index (such as the time difference between data generation and collection) are mapped to the interval [0, 1] through normalization processing to obtain the integrity quantization value, the clarity quantization value and the timeliness quantization value of the requirement data of the corresponding modality.
[0038] S14: Based on the multi-modal sample data labeled with quality labels, the initial three-dimensional quality score model is optimized and verified to obtain the final three-dimensional quality score model.
[0039] Specifically, the quality label of the multi-modal sample data is excellent, qualified or poor, which is determined by experts in the fields of medical sales and AI data processing in combination with actual scenes. When the initial three-dimensional quality score model is optimized, the multi-modal sample data is input into the initial three-dimensional quality score model to obtain the predicted quality score of each modality sample data, and the predicted quality label is determined according to the predicted quality score, and the matching degree of the predicted quality label and the real quality label is taken as the optimization target, and the gradient descent method is used to optimize the weight of the initial three-dimensional quality score model. Iterative adjustment is performed. The scoring accuracy of the initial three-dimensional quality scoring model is calculated in real time during the iteration process (accuracy = number of samples with consistent predicted quality label and true quality label / total number of samples x 100%), and the optimization is stopped when the accuracy is greater than or equal to 90%, but is not limited to the accuracy being greater than or equal to 90%; the initial three-dimensional quality scoring model verified in this way is the final three-dimensional quality scoring model, which can be directly used for quality scoring of multi-modal demand data.
[0040] S2: Based on the data quality score of each modality, a differentiated enhancement strategy is adopted for the demand data of the corresponding modality, and an enhanced multi-modal feature is output.
[0041] Further, based on the data quality score of each modality, a differentiated enhancement strategy is adopted for the demand data of the corresponding modality, and an enhanced multi-modal feature is output. The sub-steps are as follows: S21: According to the data quality score of each modality, a preset quality level and enhancement strategy correspondence table is traversed to determine the quality level corresponding to the demand data of each modality and the matching enhancement strategy; the quality level and enhancement strategy correspondence table sets rules for multiple modalities respectively, and each modality corresponds to three quality level range values, one quality level range value corresponds to one quality level, and one quality level corresponds to one enhancement strategy.
[0042] Specifically, the multiple modalities at least include: a text modality, a voice modality, an image modality and a sales scene environment signal modality. The quality level and enhancement strategy correspondence table sets rules for the text modality, the voice modality, the image modality and the sales scene environment signal modality respectively, and the specific rules are as follows: (1) The data quality score of the text modality is greater than or equal to A1, which corresponds to high-level text data and matches a lightweight enhanced text strategy; A2 is less than or equal to the data quality score of the text modality and less than A1, which corresponds to medium-level text data and matches a medium enhanced text strategy; the data quality score of the text modality is less than A2, which corresponds to low-level text data and matches a deep enhanced text strategy.
[0043] Wherein, the specific values of A1 and A2 can be flexibly adjusted according to the characteristics of the corresponding modality, and the application preferably: A1 = 0.8, A2 = 0.5.
[0044] The specific content of the lightweight enhanced text strategy, the medium enhanced text strategy, and the deep enhanced text strategy can be flexibly adjusted according to the characteristics of the corresponding modalities. Preferably, the lightweight enhanced text strategy is: lightweight deduplication and noise reduction are adopted to obtain enhanced text features; the medium enhanced text strategy is: semantic completion through a Bidirectional Encoder Representations from Transformers (BERT) model is adopted to obtain enhanced text features; and the deep enhanced text strategy is: a generative adversarial network (GAN) is adopted to generate complete semantic text to obtain enhanced text features.
[0045] (2) If the data quality score of the voice modality is greater than or equal to B1, the corresponding high-level voice data is matched with the lightweight enhanced voice strategy; if the data quality score of the voice modality is greater than or equal to B2 and less than B1, the corresponding medium-level voice data is matched with the medium enhanced voice strategy; and if the data quality score of the voice modality is less than B2, the corresponding low-level voice data is matched with the deep enhanced voice strategy.
[0046] Preferably, B1=0.8 and B2=0.5.
[0047] The specific content of the lightweight enhanced voice strategy, the medium enhanced voice strategy, and the deep enhanced voice strategy can be flexibly adjusted according to the characteristics of the corresponding modalities. Preferably, the lightweight enhanced voice strategy is: spectral subtraction noise reduction is adopted to obtain enhanced voice features; the medium enhanced voice strategy is: Mel-Frequency Cepstral Coefficients (MFCC) feature enhancement is adopted to obtain enhanced voice features; and the deep enhanced voice strategy is: a speech synthesis technology is combined to repair missing fragments to obtain enhanced voice features.
[0048] (3) If the data quality score of the image modality is greater than or equal to C1, the corresponding high-level image data is matched with the lightweight enhanced image strategy; if the data quality score of the image modality is greater than or equal to C2 and less than C1, the corresponding medium-level image data is matched with the medium enhanced image strategy; and if the data quality score of the image modality is less than C2, the corresponding low-level image data is matched with the deep enhanced image strategy.
[0049] Preferably, C1=0.8 and C2=0.5.
[0050] The specific content of the light enhancement image strategy, the medium enhancement image strategy and the deep enhancement image strategy can be flexibly adjusted according to the characteristics of the corresponding modal, and the application preferably: the light enhancement image strategy is: adaptive histogram equalization optimization is adopted to obtain enhanced image features; the medium enhancement image strategy is: the resolution is improved through super-resolution reconstruction to obtain enhanced image features; and the deep enhancement image strategy is: image repair is performed by using a GAN network (generative adversarial network) to obtain enhanced image features.
[0051] (4) If the data quality score of the sales scene environment signal modal is greater than or equal to D1, the corresponding high-level sales scene environment signal data is matched with the light enhancement sales scene environment signal strategy; if the data quality score of the sales scene environment signal modal is less than D1 and greater than or equal to D2, the corresponding medium-level sales scene environment signal data is matched with the medium enhancement sales scene environment signal strategy; and if the data quality score of the sales scene environment signal modal is less than D2, the corresponding low-level sales scene environment signal data is matched with the deep enhancement sales scene environment signal strategy.
[0052] The specific values of D1 and D2 can be flexibly adjusted according to the characteristics of the corresponding modal, and the application preferably: D1=0.8 and D2=0.5.
[0053] The specific content of the light enhancement sales scene environment signal strategy, the medium enhancement sales scene environment signal strategy and the deep enhancement sales scene environment signal strategy can be flexibly adjusted according to the characteristics of the corresponding modal, and the application preferably: the light enhancement sales scene environment signal strategy is: only extreme outliers are filtered to obtain enhanced sales scene environment signal features; the medium enhancement sales scene environment signal strategy is: a sliding average filter is adopted to obtain enhanced sales scene environment signal features; and the deep enhancement sales scene environment signal strategy is: signal completion is performed in combination with scene features to obtain enhanced sales scene environment signal features.
[0054] S22: Based on the enhancement strategy matched by the demand data of each modal, the demand data of the corresponding modal is enhanced to obtain enhanced features of the demand data of each modal, and the enhanced features of the demand data of each modal at least include enhanced text features, enhanced voice features, enhanced image features and enhanced sales scene environment signal features.
[0055] S23: The enhanced features of the demand data of each modal are cross-modal aligned and structuredly fused according to the semantic correlation degree of the medical sales scene to obtain structured multi-modal fusion features.
[0056] Specifically, the Z-score standardization method is adopted to map the enhanced text features, enhanced speech features, enhanced image features and enhanced sales scene environment signal features to the same dimensional space, so as to eliminate the dimensional differences between modalities; based on the attention mechanism, the semantic correlation weight (such as the correlation degree of the text keyword and the speech semantic label, the adaptation degree of the image key region and the environment signal intensity) between the enhanced features of the demand data of each modality is mined, and the range of the semantic correlation weight is [0, 1]; the enhanced features of the demand data of each modality are weighted and fused according to the semantic correlation weight, to form a structured multi-modal fusion feature which has both the core information of each modality and the scene adaptability.
[0057] S24: The structured multi-modal fusion feature is subjected to dual verification of feature integrity and scene adaptability, wherein the feature integrity verification is to statistically verify the coverage of the key information of each modality in the structured multi-modal fusion feature, and if the coverage is greater than or equal to a coverage threshold, it is determined that the verification is passed; the scene adaptability verification is to match the structured multi-modal fusion feature with the typical demand features of the medical sales scene, and if the matching degree is greater than or equal to a matching degree threshold, it is determined that the verification is passed; if both the dual verifications are passed, the structured multi-modal fusion feature is output as the final enhanced multi-modal feature, and if any verification is not passed, S21 is re-executed.
[0058] Specifically, the structured multi-modal fusion feature and the typical demand features of the medical sales scene are drug information integrity, customer demand recognition, scene environment correlation, etc. The specific values of the coverage threshold and the matching degree threshold are adjusted according to actual needs, and the application preferably has: coverage threshold = 90%, matching degree threshold = 95%.
[0059] S3: The enhanced multi-modal feature and the data quality score of the corresponding modality are jointly input into a pre-set spatio-temporal attention fusion model, and the spatio-temporal attention fusion model outputs a multi-modal fusion feature; then the multi-modal fusion feature is input into a pre-set deep learning model, and the deep learning model outputs a demand intention category, a priority and an urgency score.
[0060] Further, the pre-set spatio-temporal attention fusion model is a traditional spatio-temporal attention fusion model or an improved spatio-temporal attention fusion model. The expression of the multi-modal fusion feature output by the improved spatio-temporal attention fusion model is: ; Wherein, is the multi-modal fusion feature; is the basic weight corresponding to the demand data of the i-th modality; is the enhanced multi-modal feature corresponding to the demand data of the i-th modality; is the basic weight corresponding to the demand data of the i-th modality; is the enhanced multi-modal feature corresponding to the demand data of the i-th modality; is the basic weight corresponding to the demand data of the i-th modality. The time attention coefficient corresponding to the demand data of each modality; For the first The spatial attention coefficient corresponding to the demand data of each modality; For the first The data quality score corresponding to the demand data for each modality; For the first Supplementary feature items corresponding to the demand data for each modality; ,in, This represents the total number of modal types.
[0061] Specifically, the Bayesian optimization algorithm is used to determine... However, it is not limited to using Bayesian optimization algorithms. Based on the first... Calculation of the matching degree between the generation time of demand data for each modality and the timeliness of demand response. Based on the first Calculation of the correlation between demand data for each modality and sales scenarios Generate through adversarial training It is used to compensate for the lack of features in low-quality modes.
[0062] The improved spatiotemporal attention fusion model proposed in this application introduces data quality scores for each modality, adds exclusive feature supplements for each modality, and uses Bayesian optimization to determine the basic weights. Compared with the traditional spatiotemporal attention model, it can not only filter out the interference of low-quality data and fill in the gaps in modal information, but also adapt to the differences in modal weights in the pharmaceutical sales scenario. It outputs more accurate, complete and business-oriented multimodal fusion features, which can effectively improve the accuracy of subsequent tasks such as demand intent classification and priority scoring.
[0063] Furthermore, the preset deep learning model adopts existing multi-task classification models, including but not limited to the Bidirectional Encoder Representations from Transformers (BERT) multi-label classification model and the Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) fusion classification model.
[0064] S4: Combining real-time inventory data, regional pharmaceutical policies, and user profile information, based on demand intent category, priority, and urgency scores, generate personalized multimodal response content adapted to the scenario.
[0065] Furthermore, by combining real-time inventory data, regional pharmaceutical policies, and user profile information, and based on demand intent categories, priorities, and urgency scores, the following sub-steps are used to generate personalized multimodal response content adapted to the scenario: S41: The extracted real-time inventory data, regional pharmaceutical policies and user profile information, as well as the integrated demand intent categories, priorities and urgency scores, are structured and cleaned to form a standardized parameter dataset.
[0066] Specifically, real-time inventory data should include at least the quantity of medicines in stock and their out-of-stock status; regional pharmaceutical policies should include at least the scope of medical insurance reimbursement and centralized procurement policies; and user profile information should include at least historical purchase records, credit ratings, and user interaction preferences.
[0067] Structured data cleaning includes at least the following: standardizing data formats, filling in missing values, and removing redundant data.
[0068] The extracted real-time inventory data, regional medical policies, and user profile information, as well as the integrated demand intent categories, priorities, and urgency scores, are structured and cleaned using existing ETL tools (Extract-Transform-Load) or Python data processing libraries (such as Pandas) to form a standardized parameter dataset, but are not limited to existing ETL tools or Python data processing libraries.
[0069] S42: Calculate the recommendation priority of each response content by using a pre-set recommendation priority model on the standardized parameter dataset. Each response content includes at least drug recommendations, policy interpretations, and drug purchase guidelines.
[0070] Furthermore, the preset recommendation priority model can be implemented using existing recommendation methods or improved recommendation methods.
[0071] Among them, the existing recommendation methods are single-preference recommendation methods, but they are not limited to single-preference recommendation methods. The expression of the single-preference recommendation method is: ; in, The first one obtained using the single preference recommendation method The recommendation priority of each response content; For the first The response content corresponds to the user's historical purchasing frequency.
[0072] Specifically, The output may be unitless or use a 0-100 scale. As an input parameter, the unit is: times / preset cycle, such as times / month or times / quarter. The preset cycle can be set according to business needs.
[0073] The expression of the improved recommendation method is: ; wherein, is the recommendation priority of the first response content obtained by using the improved recommendation method; is the dynamic weight of the first response content, ; is the demand matching degree of the first response content corresponding to the medicine; is the policy coefficient corresponding to the first response content; is the user historical purchase frequency corresponding to the first response content; is the user credit coefficient corresponding to the first response content; is the medicine inventory sufficiency degree corresponding to the first response content; is the emergency degree coefficient corresponding to the first response content; is the medicine sales heat degree corresponding to the first response content.
[0074] Specifically, the specific value of is adjusted in real time with the response content, scene demand through reinforcement learning. Based on the demand intention category in the standardized parameter data set, is calculated, which can be realized by using the existing rule matching algorithm, but is not limited to using the existing rule matching algorithm. Based on the regional medicine policy in the standardized parameter data set, is calculated, which can be realized by using the existing policy rule mapping logic, but is not limited to the existing policy rule mapping logic. Based on the user portrait information extracted from the standardized parameter data set, is calculated, which can be realized by using the existing statistical tool, but is not limited to using the existing statistical tool. Based on the user credit level extracted from the user portrait information, is calculated, which can be realized by using the existing score mapping rule, but is not limited to using the existing score mapping rule. Based on the real-time inventory data in the standardized parameter data set, is calculated, which can be realized by using the existing inventory statistical logic, but is not limited to using the existing inventory statistical logic. Based on the emergency degree score conversion in the standardized parameter data set, is calculated, which can be realized by using the existing score conversion rule, but is not limited to using the existing score conversion rule. Based on the real-time sales data in the standardized parameter data set, The existing heat calculation algorithm can be used to realize the improved recommendation method, but is not limited to the existing heat calculation algorithm.
[0075] The improved recommendation method breaks through the limitation of the existing single preference recommendation method which only depends on the historical purchase frequency, fuses multi-dimensional data, and adapts to the scene through dynamic weight, so that the recommendation priority is more in line with the actual needs of the medicine scene, the accuracy is higher, and the mature technology is realized, and the practicability and expansibility are considered.
[0076] S43: According to the user interaction preference in the user portrait information, based on the recommendation priority and the standardized parameter data set, the personalized multi-modal response content adapted to the scene is generated.
[0077] Further, according to the user interaction preference in the user portrait information, based on the recommendation priority and the standardized parameter data set, the sub-step of generating the personalized multi-modal response content adapted to the scene is as follows: S431: Based on the recommendation priority, the response contents are sorted in descending order to generate a response content display sequence, and the response content with high recommendation priority is preferentially displayed.
[0078] Specifically, if the recommendation priority of the medicine recommendation is higher than that of the policy interpretation and the purchase guide, the medicine recommendation is preferentially displayed as the core response content, and the policy interpretation and the purchase guide are arranged in order according to the recommendation priority.
[0079] S432: According to the priority sorting result of the response content display sequence, the response content corresponding to the first N recommendation priorities is selected as the target response content, and N is a preset positive integer.
[0080] Specifically, the Top-N selection rule is adopted, and the value of N can be adjusted according to the actual scene, and the application preferably selects N=2; for example, the medicine recommendation (first priority) and the policy interpretation (second priority) are selected as the target response content.
[0081] S433: The existing multi-modal generation tool is called to determine the target presentation mode of the target response content according to the user interaction preference in the user portrait information, wherein the user interaction preference at least includes one or more of the text presentation mode, the voice presentation mode and the graphic presentation mode.
[0082] Specifically, the specific type of the existing multi-modal generation tool can be set according to the actual demand, and the application preferably selects the existing multi-modal generation tool at least including a text generation interface, a voice synthesis SDK (voice synthesis software development kit) and a graphic layout tool.
[0083] S434: Retrieve the corresponding response content template according to the target response content, fill in the real-time inventory data in the standardized parameter data set, regional medical policy, user portrait information, demand intention category, priority and emergency degree score into the response content template according to the format requirements of the target presentation mode, and obtain the initial multi-modal response content.
[0084] For example: text presentation mode: present in the form of structured list (including drug name, inventory status, medical insurance adaptation, and drug purchase process steps); voice presentation mode: convert text content into natural speech (adapt to user's preferred speech speed, such as default medium speed), and highlight the key information of high priority content; graphic presentation mode: present in the form of graphic combination (drug recommendation inventory sufficient / shortage identification, policy interpretation with medical insurance reimbursement ratio chart).
[0085] S435: Through the existing data verification tool, compare and verify the filled-in information in the initial multi-modal response content with the standardized parameter data set, ensure that the filled-in information is unbiased, and output the initial multi-modal response content as the personalized multi-modal response content adapted to the scene.
[0086] Specifically, the existing data verification tool can use data comparison software, interface verification SDK (interface verification software development kit), rule engine, database trigger, field verification script, and data consistency verification platform, but is not limited to the above types.
[0087] S5: Based on the collected user feedback data on the personalized multi-modal response content adapted to the scene and the real-time sales data, optimize the system core model to obtain an optimized system core model; the system core model at least includes a spatio-temporal attention fusion model and a deep learning model.
[0088] Further, based on the collected user feedback data on the personalized multi-modal response content adapted to the scene and the real-time sales data, the sub-steps of optimizing the system core model are as follows: S51: Collect user feedback data on multi-modal response content and real-time sales data, and the feedback data at least includes user satisfaction score and correction suggestion, and the real-time sales data at least includes recommendation conversion rate and inventory turnover efficiency.
[0089] Specifically, the existing data collection tool can be used to obtain user feedback data on multi-modal response content and real-time sales data, such as user evaluation module and sales data statistical system.
[0090] S52: Input the feedback data and real-time sales data into the preset comprehensive optimization model, and output the comprehensive optimization coefficient from the comprehensive optimization model.
[0091] The expression of the comprehensive optimization coefficient output by the comprehensive optimization model is: ; wherein, is a comprehensive optimization coefficient, ; is a dimension data weight coefficient, ; is a normalized value of the user satisfaction score; is a normalized value of the recommendation conversion rate; is a normalized value of the inventory turnover efficiency.
[0092] Specifically, the specific value of the comprehensive optimization coefficient is set according to actual needs, and the application preferably: .
[0093] S53: The comprehensive optimization coefficient is analyzed using a preset two-stage optimization threshold. When the comprehensive optimization coefficient is greater than or equal to the two-stage optimization threshold, the current parameters and structure of the system core model are maintained unchanged, and feedback data and real-time sales link data are continuously collected to dynamically monitor system performance. When the comprehensive optimization coefficient is less than the two-stage optimization threshold, two-stage optimization is triggered to obtain an optimized system core model. The two-stage optimization includes parameter-level optimization and structure-level optimization.
[0094] The parameter-level optimization includes updating the parameters in the spatio-temporal attention fusion model.
[0095] Specifically, the parameters in the spatio-temporal attention fusion model are updated , , , which can be implemented using existing machine learning frameworks (such as TensorFlow, PyTorch) gradient descent, Adam, and other optimization algorithms, so further description is omitted.
[0096] Further, the parameter-level optimization also includes updating the parameters in the recommendation priority model.
[0097] Specifically, the parameters in the improved recommendation method are updated , which can be implemented using existing machine learning frameworks (such as TensorFlow, PyTorch) gradient descent, Adam, and other optimization algorithms, so further description is omitted.
[0098] The structure-level optimization includes adjusting the layer structure of the deep learning model through neural architecture search (NAS) to improve the adaptability of the deep learning model to new scenarios.
[0099] Specifically, NAS is a mature technology for optimizing existing model structures, so further description is omitted. The specific value of the two-stage optimization threshold is set according to the actual scenario, and the application preferably is 0.7.
[0100] The steps S51-S53 of the present application break through the limitations of the prior art single feedback dimension and fixed parameter structure, and through double-dimension feedback data collection, customized comprehensive optimization coefficient calculation and double-stage optimization triggered by threshold, realize dynamic iteration of system core model parameters and structure in the medicine sales demand response system based on the multi-modal AI intelligent system, which not only improves the flexibility of the system adapting to the medicine sales scene, but also guarantees the pertinence and full-link closed-loop of the optimization.
[0101] The beneficial effects realized by the present application are as follows: (1) The medicine sales demand response method and system based on the multi-modal AI intelligent system of the present application can significantly improve the multi-modal fusion precision and anti-interference ability: the quality of the demand data of each modality is quantitatively evaluated by the three-dimensional quality scoring model, the low-quality data is optimized by the differential enhancement strategy, and the deep semantic fusion of multi-modal features is realized by the spatio-temporal attention fusion model, effectively avoiding the interference of low-quality data and greatly improving the stability and accuracy of intent recognition.
[0102] (2) The medicine sales demand response method and system based on the multi-modal AI intelligent system of the present application can comprehensively optimize personalized response and scene adaptability: the core information such as real-time inventory data, regional medicine policy and user portrait information is deeply integrated, the multi-modal response form is determined based on the recommendation priority calculation and user interaction preference, so that the response content not only fits the real-time sales scene dynamics, but also accurately matches the personalized needs of users, significantly improving the practicality of the response content and user experience.
[0103] (3) The medicine sales demand response method and system based on the multi-modal AI intelligent system of the present application has a sustained adaptation capability: by collecting feedback data of users on personalized multi-modal response content of adaptive scenes and real-time data of sales links, a double-stage optimization mechanism of the system core model is constructed, dynamic adjustment of model parameters and structure is realized, and it is ensured that the response accuracy can continuously adapt to changes in market policies, inventory status and user needs in the long-term use process, thereby prolonging the effective service period of the system.
[0104] Although the preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the scope of protection of the present application is intended to include the preferred embodiments and all changes and modifications falling within the scope of the present application. Obviously, those skilled in the art can make various modifications and changes to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and changes of the present application fall within the scope of the present application and its equivalent technology, the present application also intends to include these modifications and changes.
Claims
1. A pharmaceutical sales demand response method based on a multimodal AI intelligent system, characterized in that, Includes the following steps: S1: Obtain multimodal demand data, input the demand data of each modality into the pre-built three-dimensional quality scoring model, and output the data quality score of the corresponding modality from the three-dimensional quality scoring model; wherein, the multimodality includes at least: text modality, voice modality, image modality and sales scenario environment signal modality; S2: Based on the data quality scores of each modality, a differentiated enhancement strategy is adopted for the required data of the corresponding modality, and enhanced multimodal features are output; S3: Input the enhanced multimodal features and the corresponding data quality scores into a pre-set spatiotemporal attention fusion model, and the spatiotemporal attention fusion model outputs multimodal fusion features; then input the multimodal fusion features into a pre-set deep learning model, and the deep learning model outputs the demand intent category, priority, and urgency scores; S4: Combining real-time inventory data, regional pharmaceutical policies, and user profile information, based on demand intent category, priority, and urgency scores, generate personalized multimodal response content adapted to the scenario; S5: Based on the collected user feedback data on personalized multimodal response content for the adapted scenario and real-time data from the sales process, optimize the core model of the system to obtain the optimized core model of the system; the core model of the system includes at least: a spatiotemporal attention fusion model and a deep learning model.
2. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 1, characterized in that, The sub-steps for pre-constructing a three-dimensional quality scoring model are as follows: S11: Combining the application requirements of multimodal demand data in the pharmaceutical sales scenario, by analyzing the impact of historical multimodal demand data on the two core links of demand identification and response generation, the initial evaluation dimensions are determined. The initial evaluation dimensions should at least include data completeness, clarity, and timeliness. S12: Based on the initial evaluation dimension, access the pre-built evaluation dimension importance judgment database and select the importance judgment data of the corresponding initial evaluation dimension as the target data; The weight coefficients of each initial evaluation dimension are determined based on the target data using the analytic hierarchy process (AHP). S13: Construct an initial three-dimensional quality scoring model based on the initial evaluation dimensions and their weighting coefficients; S14: Based on multimodal sample data labeled with quality tags, optimize and validate the initial three-dimensional quality scoring model to obtain the final three-dimensional quality scoring model.
3. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 2, characterized in that, The initial three-dimensional quality scoring model includes a quantization processing layer and a quality scoring layer. The quantization processing layer receives the demand data for each modality and quantifies the completeness, clarity, and timeliness indicators of the demand data for each modality to obtain the quantified values of completeness, clarity, and timeliness of the demand data for each modality. These values are then input into the quality scoring layer, which outputs the data quality score corresponding to the demand data for each modality.
4. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 3, characterized in that, The expression for the data quality score corresponding to the demand data of each modality output by the quality scoring layer is: ; in, For the first The data quality score corresponding to the demand data for each modality; For integrity weight, For clarity weight, As a weight for timeliness, ; For the first The completeness quantification value of the demand data for each modality; For the first Clear quantifications of demand data for each modality; For the first The timeliness quantification value of demand data for each modality.
5. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 1, characterized in that, Based on the data quality scores for each modality, the following sub-steps are used to output enhanced multimodal features by employing differentiated enhancement strategies for the required data of the corresponding modality: S21: Based on the data quality scores of each modality, traverse the preset quality level and enhancement strategy correspondence table to determine the quality level and matching enhancement strategy corresponding to the required data of each modality; The quality level and enhancement strategy correspondence table sets rules for multiple modalities. Each modality corresponds to three quality level range values. Each quality level range value uniquely corresponds to one quality level, and each quality level uniquely corresponds to one enhancement strategy. S22: Based on the enhancement strategy of matching demand data of each modality, the demand data of the corresponding modality is enhanced to obtain the enhanced features of demand data of each modality. The enhanced features of demand data of each modality include at least enhanced text features, enhanced speech features, enhanced image features and enhanced sales scenario environmental signal features. S23: The enhanced features of the demand data of each modality are aligned and structurally fused across modalities according to the semantic relevance of the pharmaceutical sales scenario to obtain structured multimodal fusion features; S24: Perform dual verification of feature integrity and scenario adaptability on the structured multimodal fusion features. Feature integrity verification involves statistically analyzing the coverage of key information of each modality in the structured multimodal fusion features. If the coverage is greater than or equal to the coverage threshold, the verification is considered successful. Scenario adaptability verification involves matching the structured multimodal fusion features with the typical demand features of the pharmaceutical sales scenario. If the matching degree is greater than or equal to the matching degree threshold, the verification is considered successful. If both verifications pass, the structured multimodal fusion features are output as the final enhanced multimodal features. If either verification fails, S21 is re-executed.
6. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 5, characterized in that, Coverage threshold = 90%, matching threshold = 95%.
7. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 1, characterized in that, The following are the sub-steps for generating personalized multimodal response content tailored to specific scenarios, based on real-time inventory data, regional pharmaceutical policies, and user profile information, and according to demand intent category, priority, and urgency scores: S41: The extracted real-time inventory data, regional pharmaceutical policies and user profile information, as well as the integrated demand intent categories, priorities and urgency scores, are structured and cleaned to form a standardized parameter dataset; S42: Calculate the recommendation priority of each response content by using a pre-set recommendation priority model on the standardized parameter dataset. Each response content includes at least drug recommendations, policy interpretations, and drug purchase guidelines. S43: Based on user interaction preferences in user profile information, and using recommendation priorities and standardized parameter datasets, generate personalized multimodal response content adapted to the scenario.
8. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 7, characterized in that, Based on user interaction preferences in user profile information, and using recommendation priorities and standardized parameter datasets, the following sub-steps are used to generate personalized multimodal response content adapted to specific scenarios: S431: Based on the recommendation priority, sort the response content in descending order to generate a response content display sequence, and prioritize displaying response content with high recommendation priority; S432: Display the priority sorting results of the response content sequence, and select the response content corresponding to the top N recommended priorities as the target response content, where N is a preset positive integer; S433: Invoke the existing multimodal generation tool and determine the target presentation modality of the target response content based on the user interaction preferences in the user profile information. The user interaction preferences include at least one or more of the following: text presentation modality, voice presentation modality, and graphic presentation modality. S434: Retrieve the corresponding response content template according to the target response content, and fill the response content template with the real-time inventory data, regional medical policies, user profile information, demand intent category, priority and urgency score in the standardized parameter dataset according to the format requirements of the target presentation modality to obtain the initial multimodal response content. S435: Using existing data verification tools, compare and verify the information entered in the initial multimodal response content with the standardized parameter dataset. After ensuring that the entered information is without deviation, output the initial multimodal response content as personalized multimodal response content adapted to the scenario.
9. The pharmaceutical sales demand response method based on a multimodal AI intelligent system according to claim 1, characterized in that, Based on the collected user feedback data on personalized multimodal responses to suitable scenarios and real-time sales data, the core system model is optimized. The sub-steps for obtaining the optimized core system model are as follows: S51: Collect user feedback data on personalized multimodal response content for the appropriate scenario and real-time sales data. Feedback data shall include at least user satisfaction ratings and correction suggestions, and real-time sales data shall include at least recommendation conversion rate and inventory turnover efficiency. S52: Input feedback data and real-time sales data into the preset comprehensive optimization model, and the comprehensive optimization model outputs the comprehensive optimization coefficient; S53: Analyze the comprehensive optimization coefficient using a preset dual-level optimization threshold. When the comprehensive optimization coefficient is greater than or equal to the dual-level optimization threshold, maintain the current parameters and structure of the system core model unchanged, and continuously collect feedback data and real-time data from the sales process to dynamically monitor system performance. When the comprehensive optimization coefficient is less than the dual-level optimization threshold, trigger dual-level optimization to obtain the optimized system core model. Dual-level optimization includes parameter-level optimization and structural-level optimization. Parameter-level optimization includes updating the parameters in the spatiotemporal attention fusion model; structural-level optimization includes adjusting the layer structure of the deep learning model through neural architecture search.
10. A pharmaceutical sales demand response system based on a multimodal AI intelligent system, characterized in that, include: Data acquisition subsystem, pharmaceutical sales demand response center, and user terminal; The data acquisition subsystem is used to collect multimodal demand data and transmit it to the pharmaceutical sales demand response center; it also collects user feedback data on personalized multimodal response content for the appropriate scenario and real-time sales data and sends it to the pharmaceutical sales demand response center. Pharmaceutical Sales Demand Response Center: Used to execute the pharmaceutical sales demand response method based on a multimodal AI intelligent system as described in any one of claims 1-9; User-side: Provides users with a multimodal request submission portal; receives and displays personalized multimodal response content adapted to specific scenarios output by the pharmaceutical sales demand response center.