Multi-modal AI intelligent agent based on deep learning and backup method

Through multimodal AI agents and backup methods based on deep learning, the efficiency bottlenecks and security problems of multimodal data processing are solved, intelligent generation and security protection of multi-source information are realized, and content quality and production efficiency are improved.

CN120597925APending Publication Date: 2025-09-05GUANGZHOU SHUYUANCHANGLIAN SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510745264.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

When processing massive and heterogeneous multimodal data, the existing technology has efficiency bottlenecks and lack of intelligence, making it difficult to achieve in-depth understanding, correlation analysis and intelligent integration of multi-source information. The digital assets lack a convenient and reliable localized protection mechanism, resulting in insufficient quality, timeliness and personalization of generated content, and there is a risk of loss or management chaos.

Method used

Multimodal AI agents and backup methods based on deep learning receive text, voice, image and video data through the multimodal cognitive center, dynamically generate e-commerce copywriting, conference minutes, design artwork and presentations, and perform encrypted and compressed backups through the asset-aware backup engine, using cross-modal association and intention generation adversarial network optimization content generation, combining blockchain verification to ensure data security.

Benefits of technology

It realizes unified processing and intelligent generation of multi-modal information, improves the accuracy of multi-source information correlation, ensures high attractiveness and timeliness of generated content, reduces the misjudgment rate, and ensures the security and integrity of data through value hierarchical compression and blockchain verification mechanisms, forming a full-link closed-loop efficient digital content production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597925A_ABST
    Figure CN120597925A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal AI intelligent agent based on deep learning and a backup method, and relates to the technical field of artificial intelligence and data storage, and the method comprises the following steps: S1, a multi-modal cognitive center receives heterogeneous input data of texts, voices, images and videos; and S2, the dynamic content generation system generates an e-commerce document, a conference summary, a design drawing manuscript and a presentation manuscript based on the output of the cognitive center. According to the multi-modal AI intelligent agent based on deep learning and the backup method, text, voice and visual data are uniformly processed through the cross-modal semantic distillation network, and the problem of information splitting is fundamentally solved. Emotional mapping of commodity video pictures and user comments is established by using a time sequence alignment attention mechanism, so that the multi-source information association precision is improved; and in combination with the dynamic optimization capability of the intention generative adversarial network, autonomous iterative updating of the e-commerce copywriting is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and data storage technology, and specifically to a multimodal AI agent and backup method based on deep learning. Background Art

[0002] In the current era of digital information explosion, users face significant efficiency bottlenecks and a lack of intelligence when processing massive, heterogeneous, and multimodal data and efficiently generating valuable derivative content based on it. Existing technologies often rely on fragmented, single-function tools to process information from different modalities or perform specific tasks. For example, e-commerce operations require manual or disparate tools to capture product information, analyze and review it, and write copy. Meeting minutes require dedicated personnel to time-consumingly organize lengthy recordings to extract key points. Promotional design requires specialized software or repeated communication. Presentation document creation involves tedious information collection, logical organization, and layout design. These processes not only consume significant manpower and time, but also, due to the fragmented nature of tools and low levels of automation, they struggle to achieve deep understanding, correlation analysis, and intelligent integration of multi-source information. This results in insufficient quality, timeliness, and personalization of the generated content. Furthermore, these valuable digital assets generated by users or rudimentary tools, as well as the raw data generated during the work process, generally lack convenient, reliable, and intelligent local protection mechanisms, exposing them to the risk of loss or mismanagement. Users urgently need an overall solution that can fundamentally and uniformly understand and process multimodal information and intelligently generate diverse application content, while ensuring that these digital achievements are safe and controllable. Summary of the Invention

[0003] (1) Technical problems solved In response to the shortcomings of the existing technology, the present invention provides a multimodal AI agent and backup method based on deep learning, which solves the problem of how to fundamentally and uniformly understand and process multimodal information and intelligently generate diversified application content, while ensuring the security and controllability of these digital achievements.

[0004] (2) Technical solution To achieve the above objectives, the present invention is implemented through the following technical solutions: a multimodal AI agent and backup method based on deep learning, comprising the following steps: Step S1: The multimodal cognitive center receives heterogeneous input data of text, voice, image and video; Step S2: The dynamic content generation system generates e-commerce copywriting, meeting minutes, design drawings, and presentations based on the output of the cognitive center; Step S3: The asset-aware backup engine performs encrypted and compressed backup of the content output by the generation system and the original input data; Step S4: the output end of the cognitive center is connected to the input end of the generation system, and the output end of the generation system is connected to the input end of the backup engine.

[0005] Preferably, the dynamic content generation system includes an e-commerce copywriting generation unit, a speech analysis unit, a design generation unit, and a PPT generation unit; The e-commerce copywriting generation unit captures product details pages, user reviews, and video introductions, and dynamically updates product titles and promotional copywriting; The speech analysis unit transcribes sales calls and meeting recordings, and outputs customer demand analysis reports and structured meeting minutes; The design generation unit generates cross-platform visual materials according to natural language instructions; The PPT generation unit automatically generates training materials, marketing plans and project report documents based on the subject description.

[0006] Preferably, the multimodal cognitive center in step S1 executes the following steps in sequence: Adaptive coding stage: extracting underlying features through text encoder, speech encoder and video encoder respectively; Cross-modal association stage: Use the temporal alignment attention mechanism to associate feature vectors of different modalities and establish a sentiment mapping relationship between product videos and user reviews; Content synthesis stage: The associated features are input into the intent generative adversarial network, the generator outputs the draft content, and the discriminator evaluates the content value based on the human cognitive model.

[0007] Preferably, the e-commerce copywriting generation unit includes: Deploy a dual-loop feedback engine, with the inner loop monitoring the click-through rate of copywriting through real-time tracking data, and the outer loop combining market sentiment analysis results; When the copywriting effect is detected to be declining, the key selling point features in the user behavior data are strengthened based on the gradient reversal layer; The generator updates the copy content online based on the enhanced features.

[0008] Preferably, the voice analysis unit performs double verification: Simultaneously extract the semantic features of the speech transcription text and the emotional features of the voiceprint of the original recording; When the difference between the voiceprint emotional feature and the text semantic confidence exceeds the preset threshold, the multi-channel recalibration module is triggered to perform secondary analysis; The output end is connected to the generation interface of the customer pain point analysis report and the meeting action item list.

[0009] Preferably, the design generation unit adopts a staged evaluation, including: Structural verification stage: Verify the component layout of the design draft through the preset typesetting rule engine; Aesthetic verification stage: Use the visual semantic matching model to calculate the consistency score of the design draft image and text; When both phases of evaluation are passed, the final design draft is sent to the backup engine.

[0010] Preferably, the asset-aware backup engine in step S3 performs encrypted compression backup on the content output by the generation system and the original input data, including: The file identification module marks the value hierarchy of original data, intermediate processed data and final output based on metadata; The core asset compression module uses a neural differential encoder to extract the semantic features of design drawings and PPT documents, achieving high-compression storage; The backup verification module is embedded in the blockchain light node, and bit-level data integrity comparison is performed through the hash chain during incremental backup.

[0011] The control logic of the asset-aware backup engine is as follows: the backup engine automatically classifies file value levels based on file metadata and uses differential compression technology for intelligently generated design drafts and documents. The compression process extracts the deep structural features and style parameters of the file. For example, the font skeleton coordinates and color gradient rules are retained in the design draft, and the topological relationship of the chart is stored in the document. The compressed feature package is encrypted and transmitted to the storage node. During incremental backup, the data integrity verification mechanism is activated, and a digital fingerprint is generated for the current file by block and compared with the historical fingerprint chain. When an abnormal change in the fingerprint of a specified block is detected, the nature of the change is automatically analyzed. If the price is modified by an authorized user, the backup is executed. If an unauthorized watermark is detected, the process is frozen and an alarm is issued. The system regularly performs full verification and restores the damaged block from the three most recent valid versions when data corruption is found.

[0012] Preferably, the process of extracting semantic features of design drawings and PPT documents using a neural differential encoder includes: The encoding end extracts the deep semantic feature vector of the file; Compress the feature vector through the feature distillation layer; The decoding end reconstructs visual elements and text content based on semantic feature vectors.

[0013] Preferably, the e-commerce copywriting generation unit, speech analysis unit, design generation unit and PPT generation unit in the dynamic content generation system constitute a federated skill network; each unit shares the feature extraction layer of the multimodal cognitive center, but has independent generator fine-tuning parameters; the units exchange model parameters through the knowledge distillation channel, where the hot word model of the e-commerce copywriting unit is converted into the creative prompt vector of the short video script unit.

[0014] The collaborative evolution of the federated skill network involves the following: each functional unit shares the core feature extraction layer but retains independent fine-tuning parameters. For example, the e-commerce unit enhances the weight of electronic product parameters, and the design unit optimizes the visual element generation strategy. The knowledge distillation channel regularly compresses the learning results between transmission units, for example, converting the efficient hot word model of the copywriting unit into a low-dimensional vector and transmitting it to the design unit. The receiving unit tests the compatibility of the new parameters in a sandbox environment and updates the local generation module after verification. The user's operational behavior on the updated content is captured as feedback signals. Positive feedback triggers the broadcast of parameters across the entire network, and frequent generation failures initiate version rollbacks. The channel maintenance process regularly optimizes the transmission protocol and restructures the transmission model with high error rates to improve accuracy.

[0015] Preferably, the knowledge distillation channel includes: Low-rank matrix decomposition is used to compress the model parameters for transmission; The parameter receiving unit updates the local generator according to the new parameters; The updated generator output is fed back to the parameter adjustment interface of the multimodal cognitive center.

[0016] Implementation safeguards for the deep learning-based multimodal AI agent and backup method include: The speech analysis unit incorporates a voiceprint and semantic conflict resolution mechanism. When a significant discrepancy between emotional expression and textual content is detected, the analysis timeframe is automatically expanded and historical recording samples are linked. The design unit employs an element decoupling strategy to handle multi-style conflicts, separating platform compliance requirements from artistic style elements and achieving compatible output through dynamic adjustments. The backup system establishes redundant storage nodes for key data blocks, initiating multi-version collaborative repair when full verification reveals damage. The skill network monitors newly connected modules during the trial period, automatically isolating them to an independent sandbox environment if the fault tolerance threshold is exceeded.

[0017] (3) Beneficial effects The present invention provides a multimodal AI agent and backup method based on deep learning. It has the following beneficial effects: (1) This deep learning-based multimodal AI agent and backup method uses a cross-modal semantic distillation network to uniformly process text, voice, and visual data, fundamentally resolving the problem of information fragmentation. A temporal alignment attention mechanism is used to establish an emotional mapping between product videos and user reviews, improving the accuracy of multi-source information correlation. Combined with the dynamic optimization capabilities of the intent-generating adversarial network, it enables autonomous iteration and updating of e-commerce copy. A dual-loop feedback engine automatically enhances core selling point features based on real-time user behavior and market sentiment, ensuring that generated content remains highly engaging. The speech analysis unit utilizes a dual verification mechanism of voiceprint and semantics. When a serious conflict between emotional expression and textual content is detected, contextual backtracking is used to accurately identify the true user intent, reducing the error rate. The design draft generation process is verified in stages based on structural specifications and aesthetic semantics, effectively balancing platform compliance and creative freedom.

[0018] (2) This deep learning-based multimodal AI agent and backup method achieves efficient knowledge transfer through a federated skill network architecture. The hot word model of the e-commerce copywriting unit is converted into a creative vector for the design unit through low-rank compression, reducing the cost of integrating new functional modules. The asset-aware backup engine, based on a value-grading strategy, performs differential neural coding compression on intelligently generated design drafts and documents, reducing storage overhead while preserving complete semantics. The blockchain verification module ensures the immutability of backup files through bit-level data comparison and a multi-version collaborative repair mechanism. The closed-loop feedback system of the knowledge distillation channel captures user operational behavior to drive parameter self-optimization, resulting in the continuous evolution of system capabilities. Ultimately, a full-link closed loop covering data understanding, content generation, and asset protection is constructed, comprehensively improving the efficiency of digital content production. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic diagram of the overall process of the present invention; Figure 2 This is a control logic timing diagram of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0021] See also Figure 1 and Figure 2 The present invention provides a technical solution: a multimodal AI agent and backup method based on deep learning, comprising the following steps: Step S1: The multimodal cognitive center receives heterogeneous input data of text, voice, image and video; Step S2: The dynamic content generation system generates e-commerce copywriting, meeting minutes, design drawings, and presentations based on the output of the cognitive center; Step S3: The asset-aware backup engine performs encrypted and compressed backup of the content output by the generation system and the original input data; Step S4: The output end of the cognitive center is connected to the input end of the generation system, and the output end of the generation system is connected to the input end of the backup engine.

[0022] It's important to note that, in its implementation, the cognitive hub first performs feature normalization after receiving webpage text, voice recordings, and product videos. A dedicated encoder extracts key entity phrases and their associated attributes from the text data. For example, it can identify noise reduction depth parameters and applicable scenario keywords from a headphone product description. The voice stream is then framed and converted into a sequence of spectral features, simultaneously isolating the speaker's voiceprint.

[0023] The video data is segmented by scene, and for each segment, the main image features and the speech transcription text are simultaneously extracted. The cross-modal association engine is then activated. When a user's voice comment is detected complaining about battery life, it automatically associates the video clip with the battery capacity. The consistency between the emotional intensity of the voice and the text description is compared. If the voice emotion is marked as strong dissatisfaction but the battery parameters are fully displayed in the video, suggestions for product description improvements are generated.

[0024] The associated features are input into the content generation network, which includes a draft generator and a quality discriminator. The discriminator simulates the evaluation standards of professional users to detect whether the draft contains excessive technical terms or insufficient emotional motivation, triggering the generator to iteratively optimize until the output is in line with human expression habits.

[0025] The e-commerce copywriting unit continuously monitors user click data for published copy. When click-through rates continuously decline beyond a set threshold, it automatically crawls social media hot topics and competitive product promotion strategies. The system identifies key selling points in historically high-converting copy and adjusts feature weights based on current public opinion trends. For example, if energy conservation is detected to be gaining popularity, the weight of energy efficiency parameters in the feature vector is increased. The generator retains proven expression templates, incorporates current hot topics, and adds scenario-based use case descriptions. The voice analysis unit performs dual-channel analysis on call recordings. The speech transcription engine outputs text content and confidence scores, and the voiceprint analysis module calculates emotional fluctuation curves. When the text indicates customer satisfaction but the voiceprint shows signs of excitement, the system automatically retrieves the call context to analyze the transition semantics. After confirming the true intent, it generates a risk-flagged report. The design draft generation process strictly implements a two-stage verification process. First, the text-to-image ratio is adjusted according to the platform design specifications. Then, the visual semantic model is used to assess the theme fit. Any non-compliant elements are replaced by matching historical case library.

[0026] The dynamic content generation system includes an e-commerce copywriting generation unit, a voice analysis unit, a design generation unit, and a PPT generation unit; The e-commerce copywriting generation unit captures product detail pages, user reviews, and video introductions, and dynamically updates product titles and promotional copywriting; The voice analysis unit transcribes sales calls and meeting recordings, outputting customer demand analysis reports and structured meeting minutes; The design generation unit generates cross-platform visual materials based on natural language instructions; The PPT generation unit automatically generates training materials, marketing plans and project report documents based on the topic description.

[0027] It should be further explained that, in the specific implementation process, the dynamic content generation system first automatically captures the target product's details page HTML source code, user evaluation JSON data and related short video introduction files through the e-commerce copy generation unit, parses the function parameters and usage scenario descriptions from the details page, extracts high-frequency emotional keywords and pain point descriptions from user reviews, and identifies voice-over key sentences and visual focus areas from the video introduction; based on the fusion feature vector output by the multimodal cognitive center, the initial copy containing the product selling points is generated in the first round, and at the same time, a tracking script is deployed to monitor the click-through rate, residence time and conversion rate data of the e-commerce platform users on the copy to form an internal loop feedback.

[0028] When it is monitored that the click-through rate has dropped by more than 15% for three consecutive days, the external circulation market public opinion analysis module is activated to crawl the promotion strategies of competing products and hot topics on social platforms, and the user behavior data and public opinion analysis results are input into the gradient reversal layer to strengthen the selling point dimensions in the feature vector that are strongly related to the current market trend. If it is detected that the popularity of public opinion related to "energy saving" is increasing, the weight factor of the energy efficiency parameter is increased to drive the generator to output updated copy with dynamic promotional language.

[0029] The speech analysis unit receives sales call recordings and conference recording streams, extracts Mel-spectrogram features through a pre-trained speech encoder for transcription, and simultaneously starts the voiceprint sentiment analysis model to calculate the emotional fluctuation curve in the recording; when "price" related semantics appear in the transcribed text with a confidence level of 90%, but the voiceprint sentiment analysis shows that the user's speaking speed is faster and the pitch is higher, that is, the characteristics of excitement, the multi-channel recalibration module is triggered to retrieve the context 5 minutes before and after the call for secondary semantic analysis, and after confirming the customer's actual concerns, a customer demand report is generated with the core pain points and priorities marked; key phrases such as "need to follow up" and "need to coordinate resources" are automatically marked in the conference recording processing, and the speaker's voiceprint ID is associated to generate an action item list with responsible persons and deadlines.

[0030] The design generation unit receives the user's natural language input for the "retro-style promotional poster", parses it, and calls the cross-modal correlation features of the multimodal cognitive center to generate a draft. The typesetting rule engine first verifies whether the ratio of text and image areas meets the platform specifications. For example, the text ratio of the Xiaohongshu cover must be less than 30%. The CLIP model is then used to calculate the semantic matching degree between the poster's visual elements and the "retro" theme. If both stages of verification pass, the final design draft is output to the backup queue.

[0031] The PPT generation unit automatically divides the "smart home market analysis" topic description submitted by the user into three logical modules: market trends, competitive product comparison, and strategic recommendations. It retrieves chart templates from historical projects from the backup engine to fill in data, and inserts the customer pain point heat map output by the voice analysis unit to generate an editable presentation.

[0032] The multimodal cognitive hub in step S1 executes the following steps in order: Adaptive coding stage: extracting underlying features through text encoder, speech encoder and video encoder respectively; Cross-modal association stage: Use the temporal alignment attention mechanism to associate feature vectors of different modalities and establish a sentiment mapping relationship between product videos and user reviews; Content synthesis stage: The associated features are input into the intent generative adversarial network, the generator outputs the draft content, and the discriminator evaluates the content value based on the human cognitive model.

[0033] It should be further explained that, in the specific implementation process, after receiving heterogeneous input data, the multimodal cognitive center first performs adaptive encoding, including: using a domain-adapted word vector encoder for the product details page text to extract 128-dimensional feature vectors of functional parameter entities and scene description phrases; generating a time-frequency graph for the user evaluation voice stream through the Mel spectrum conversion layer, and outputting the emotion polarity intensity sequence through the pre-trained voice encoder; sampling key frames at intervals of 1 second for the product video, and using a lightweight visual encoder to extract the product main area features in the picture and the OCR recognition text features.

[0034] Then, the cross-modal association stage is started, and the text feature vector, speech emotion sequence, and video key frame features are input into the time-aligned attention mechanism. When the user's evaluation voice contains "short battery life" and the confidence level exceeds 80%, the system automatically associates the video clips showing the battery capacity, such as the battery percentage icon in the close-up shot. The semantic similarity between the speech emotion intensity and the video text description is calculated to establish a mapping relationship between "battery life complaints-insufficient battery parameter display".

[0035] After completing feature association, the content synthesis stage begins, and the fused feature vector is input into the intent-generating adversarial network (IGAN). The generator first generates a draft of the copy containing a comparison of technical parameters. The discriminator simulates the cognitive patterns of human professional buyers to evaluate the value of the copy, including: detecting whether there are excessive technical terms, whether the order of key selling points conforms to reading habits, and whether the density of emotional motivational words meets the standards. When the discriminator detects that the proportion of technical parameters in the draft exceeds 60% and the density of emotional words is lower than the threshold, it feeds back a weight adjustment signal to the generator, triggering the regeneration of a version with scenario-based language.

[0036] For meeting recording processing, the speaker's voiceprint ID is bound to the decision-making keywords in the transcribed text during the feature association stage. Decision-making keywords include "approval" and "rejection." If the same speaker repeatedly uses the semantics of "insufficient budget" within 10 minutes and the voiceprint fluctuates significantly, this feature is marked as a high-priority pain point and input into IGAN, ultimately generating meeting minutes with risk warning indicators.

[0037] The e-commerce copywriting generation unit includes: Deploy a dual-loop feedback engine, with the inner loop monitoring the click-through rate of copywriting through real-time tracking data, and the outer loop combining market sentiment analysis results; When the copywriting effect is detected to be declining, the key selling point features in the user behavior data are strengthened based on the gradient reversal layer; The generator updates the copy content online based on the enhanced features.

[0038] It should be further explained that, in the specific implementation process, after the e-commerce copy generation unit deploys the dual-loop feedback engine, the inner loop captures the user's interactive behavior on the current copy in real time through the embedding script, including: when it is monitored that the click-through rate of a certain headphone copy drops by 40% within 24 hours but the number of visits to the product details page is stable, it is determined that the copy is not attractive enough, and the outer loop market public opinion analysis module is automatically triggered; this module calls the social media API to capture the promotional data and hot search topics of competing products in the past 72 hours, and recognizes that the discussion volume related to "active noise reduction" has increased by 120% and that the titles of competing products all highlight this feature.

[0039] After receiving the user behavior data from the inner loop and the public opinion analysis results from the outer loop, the gradient reversal layer first locates the paragraph with the highest conversion rate in the historical version of the copy, which contains the description "noise reduction depth of up to 35dB". However, the current copy overemphasizes "lossless sound quality", resulting in a low weight for this selling point. Therefore, the "active noise reduction" related dimension in the feature vector is strengthened by 3 times the weight, while suppressing the "material craftsmanship" feature that is weakly correlated with the current public opinion.

[0040] The generator performs online fine-tuning based on the enhanced feature vector, including: retaining the "immersive experience" speech template in high-conversion rate evaluation, inserting the hot public opinion keyword "transparency mode", and adding a scenario-based description of "actual measurement of the sound insulation effect of subway commuting"; continuously monitoring after outputting the updated copy. If the click-through rate of the new copy rises above the baseline value within 48 hours after the release of the new copy and does not cause an increase in the bounce rate, the weight adjustment strategy will be stored in the optimization knowledge base.

[0041] When optimizing the copywriting of similar products in the future, the verified weight enhancement scheme in the knowledge base will be used first. If the click-through rate does not increase by 5% after two consecutive optimizations, the gradient reversal layer parameters will be reset and full feature analysis will be started.

[0042] The Voice Analysis Unit performs two-factor authentication: Simultaneously extract the semantic features of the speech transcription text and the emotional features of the voiceprint of the original recording; When the difference between the voiceprint emotional feature and the text semantic confidence exceeds the preset threshold, the multi-channel recalibration module is triggered to perform secondary analysis; The output end is connected to the generation interface of the customer pain point analysis report and the meeting action item list.

[0043] It should be further explained that, during the specific implementation process, when the speech analysis unit processes the sales call recording, it simultaneously performs speech transcription and voiceprint emotion analysis, including: the transcription engine divides the audio stream into 5-second segments, outputs text through the pre-trained speech model and marks the confidence of each sentence; the voiceprint analysis module extracts the fundamental frequency jitter rate, the sudden change in speaking speed and the energy peak, and when the fundamental frequency jitter rate exceeds 0.25 and the speaking speed suddenly increases by 40%, it is marked as "excited emotion".

[0044] If the word "refund" appears in the transcribed text with a confidence level of 95%, but the voiceprint analysis shows the emotion is marked as "calm," where "calm" refers to an emotion with a fundamental frequency jitter rate of less than 0.1, the semantics and emotion are determined to match, and a regular priority customer feedback report is generated. Conversely, if the text shows "satisfied" with a confidence level of 90% but "excitement" is detected, the multi-channel recalibration module is immediately triggered. Excitement refers to an emotion with a sudden increase in speech speed by half and an energy peak exceeding the threshold. The process of immediately triggering the multi-channel recalibration module includes: retrieving the 2-minute context before and after the call, detecting whether there is real intention guided by transition words such as "but" and "actually", and analyzing the emotional patterns of similar voiceprint IDs in other calls in the same batch.

[0045] If the secondary analysis confirms that the customer actually expressed a complaint, such as "three repairs failed to resolve the problem" appears in the context, a high-priority pain point report will be output and a red warning icon will be marked; for meeting recordings, if speaker A in the transcribed text proposes "the budget needs to be increased" and the voiceprint emotion is stable, but speaker B responds "impossible" within 10 seconds with a violent fluctuation in the voiceprint, the conversation between the two will be automatically linked to generate an action item "Budget dispute - requires high-level arbitration" and pushed to the project manager's to-do list.

[0046] The design generation unit adopts a phased assessment: Structural verification stage: Verify the component layout of the design draft through the preset typesetting rule engine; Aesthetic verification stage: Use the visual semantic matching model to calculate the consistency score of the design draft image and text; When both phases of evaluation are passed, the final design draft is sent to the backup engine.

[0047] It should be further explained that, in the specific implementation process, after the design generation unit receives the user's "Summer Beverage Promotion Xiaohongshu Cover" command, it first generates a draft containing beverage pictures, promotional slogans and price information, and enters the structure verification stage, including: calling the Xiaohongshu platform design specification library to detect whether the text area exceeds 30% of the total canvas area, whether the product body is centered and not blocked by text, and whether the brand LOGO size meets the minimum pixel requirements; if it is detected that the font of the promotional slogan is too large, resulting in an excessive proportion of text, the font size will be automatically reduced and the layout will be adjusted to a compliant state.

[0048] Manuscripts that have passed the structural check flow into the aesthetic verification stage, where the visual elements in the draft are extracted and semantic matching is calculated with the user command keywords "summer" and "refreshing". When the CLIP model outputs a similarity score lower than 0.7, it is determined that the image and text do not match the theme. The visual elements include images of mango slices and water drop splashing effects. At this time, the historical successful case library is searched and it is found that the highly matching designs all use cool backgrounds and dynamic fruit materials. Therefore, the background of the draft is automatically replaced with mint green, and a frame animation of the mango peeling process is added to increase the matching degree to above 0.85.

[0049] If the user adds an "Instagram style" requirement but the system detects that the style conflicts with platform specifications, such as the thin fonts commonly used in the Instagram style, which may lead to poor readability on mobile devices, the system will activate the balanced mode, including: retaining the white space ratio and simple composition of the Instagram style, but bolding the font to the standard threshold, to generate a final draft that meets the platform requirements and retains the core style elements.

[0050] A final compliance review will be conducted before the final draft is output. Only after confirming that the text accounts for 28%, the LOGO pixels meet the standards, and the style matching degree is 0.82, will it be transferred to the backup queue.

[0051] The asset-aware backup engine in step S3 performs encrypted and compressed backup of the generated system output and the original input data, including: The file identification module marks the value hierarchy of original data, intermediate processed data and final output based on metadata; The core asset compression module uses a neural differential encoder to extract the semantic features of design drawings and PPT documents, achieving high-compression storage; The backup verification module is embedded in the blockchain light node, and bit-level data integrity comparison is performed through the hash chain during incremental backup.

[0052] It should be further explained that, during the specific implementation process, when the asset-aware backup engine receives the final draft file output by the design generation unit, the file identification module first parses its metadata, including: if it is detected that the file creator is "AI Designer v3", the modification timestamp is within 2 hours and it is marked as "final draft", it will be classified as a high-value core asset; for meeting minutes files, when the metadata shows that it contains a "red alert" label or is associated with more than 3 action items, the backup priority will be automatically increased to the highest level.

[0053] When the core asset compression module processes the Xiaohongshu cover design draft, the neural differential encoder extracts deep semantic features from the poster, including the HSV value of the main color, the text layout topology structure, and the product body outline vector. Through the feature distillation layer, the original file's 15MB PSD layered data is compressed into a 1.8MB semantic feature package, retaining the cool color gradient parameters and font skeleton information.

[0054] The backup verification module starts the blockchain light node during the incremental backup process, including: when the user modifies the promotional price in the design draft, the SHA-256 hash value is calculated for the new version and compared with the hash chain of the previous backup at the bit level; if the binary data in the price area is detected to have changed but the brand LOGO block has not changed, a "legal modification" mark is generated to execute the backup; but if an abnormal change is found in the LOGO block hash, such as the implantation of an unauthorized watermark, the backup process is immediately frozen and a tampering alert is sent.

[0055] For PPT project report documents, the backup engine regularly initiates full verification at 20:00 every Friday, comparing the local file with the hash chain of the cloud backup node on a 128-bit block-by-block basis. When the difference rate of three consecutive blocks exceeds 5%, the fragment repair mode is triggered, and the damaged block is restored from the three most recent valid versions.

[0056] The neural differential encoder is used to extract semantic features of design drawings and PPT documents, including: The encoding end extracts the deep semantic feature vector of the file; Compress the feature vector to 10% to 15% of its original size through the feature distillation layer; The decoding end reconstructs visual elements and text content based on semantic feature vectors.

[0057] It should be further explained that in the specific implementation process, when the neural differential encoder processes the cover design draft of Xiaohongshu, the encoding end first separates the text layer and the image layer, including: extracting the font skeleton topology structure of the promotional slogan "Second piece half price", including the stroke connection point coordinate sequence and tilt angle value, and extracting the main color gradient parameters and contour key point vector coordinates in the HSV color space for the mango main image, where the main color gradient parameter is the H value gradient from 35° to 50°.

[0058] The feature distillation layer fuses the text skeleton data and the image contour coordinates into a unified semantic feature vector. By removing the redundant coordinate differences between adjacent points, it compresses it to 12% of the original data volume and retains the minimum reversible unit. For example, at least 3 coordinates are retained for the turning points of the text.

[0059] When processing a PPT project report document, the encoding end identifies the smart home concept diagram on the title page, extracts the device connection relationship topology and the color coding rules of the data flow diagram, and the distillation layer converts the topological relationship into a sparse representation of the connection matrix. After compression, only the positions of non-zero elements and the color mapping table are stored. Among them, the device connection relationship topology includes hierarchical arrows pointing from "router" to "smart speaker" in one direction and from "smart speaker" to "lamp" in one direction; the color coding rules of the data flow diagram include blue for user behavior and red for system response.

[0060] When the decoding end reconstructs the design draft, it restores the tilted font effect of the "half-price" promotional slogan through Bezier curve fitting based on the skeleton topology data in the semantic feature vector, and then fills the orange and yellow color of the mango image according to the HSV gradient parameters; for PPT documents, the device topology diagram is restored based on the sparse data of the connection matrix. If it is detected that the color code of a node is missing, such as the undefined color of the lamp node, it automatically associates the blue code of the parent node router and adds a transparency difference mark.

[0061] After the reconstruction is completed, reversibility verification is performed, including: pixel-level hash comparison of the reconstructed file and the original file in key areas. If the difference rate exceeds 2%, the compensation mechanism is activated, that is, the most recent valid feature package is retrieved from the backup engine for local coverage to ensure that the visual consistency is met before transmission to the storage node. Among them, the key areas include price numbers and brand logos.

[0062] The e-commerce copywriting generation unit, speech analysis unit, design generation unit, and PPT generation unit in the dynamic content generation system constitute a federated skill network. Each unit shares the feature extraction layer of the multimodal cognitive hub, but has independent generator fine-tuning parameters. Model parameters are exchanged between units through the knowledge distillation channel, where the hot word model of the e-commerce copywriting unit is converted into the creative prompt vector of the short video script unit.

[0063] It should be further explained that, in the specific implementation process, when the federated skill network is initialized, the e-commerce copy generation unit loads special fine-tuning parameters for the 3C category, such as the weight enhancement coefficient of "processor model" and "battery life parameters", and the speech analysis unit activates the medical industry terminology library, such as the semantic recognition priority of "clinical trial" and "dose adjustment". When each unit runs independently, it calls the underlying feature extraction layer of the multimodal cognitive center through a unified interface, that is: when the short video script unit needs to generate "smartphone review video copy", the shared layer extracts the chip model frame image and running score data chart features from the product video, but the script unit's own fine-tuning parameters suppress the density of professional terms to below 30%.

[0064] The knowledge distillation channel starts synchronization every 24 hours, including: the e-commerce copywriting unit converts the recent high-frequency conversion hot word model into a 512-dimensional prompt vector, which is compressed into a 32-dimensional core vector through low-rank matrix decomposition and transmitted to the short video script unit. Among them, the high-frequency conversion hot word model includes "AI photography", which increases the click-through rate of headphone copy by 120%; after the script unit receives it, if a "photography" related theme task is detected, the core vector is integrated with the local parameters to generate creative instructions, such as "show the night portrait light spot effect", and the completion rate data of the newly generated script is fed back to the distillation channel.

[0065] When a financial risk control module is added to the voice analysis unit, a dedicated parameter set is automatically registered with the network, such as the voiceprint emotion thresholds for "credit default" and "risk assessment". The "high-risk" related hot word model in the e-commerce unit is obtained through the distillation channel, and the telemarketing risk warning script is generated through transfer learning and adaptation to the financial scenario. If the false alarm rate of the new module exceeds 5% within 7 days of its launch, the version will be rolled back to the previous stable parameters, and the knowledge output of the module will be frozen until the false alarm rate meets the standard.

[0066] The knowledge distillation pipeline includes: Low-rank matrix decomposition is used to compress the model parameters for transmission; The parameter receiving unit updates the local generator according to the new parameters; The updated generator output is fed back to the parameter adjustment interface of the multimodal cognitive center.

[0067] It should be further explained that in the specific implementation process, when the knowledge distillation channel performs parameter transmission, the e-commerce copywriting unit reduces the dimensionality of the high-weight features of wireless fast charging in the hot word model through low-rank matrix decomposition, including: the original 512-dimensional parameter matrix is ​​decomposed into the product of a 16-by-16 basis matrix and a 32-dimensional coefficient vector, with a compression ratio of 3:1, and is transmitted to the PPT generation unit after encryption.

[0068] After parsing the parameter package, the receiving unit first creates a sandbox environment in the local generator to test compatibility: when generating a Bluetooth headset marketing PPT, the wireless fast charging coefficient vector is integrated into the theme template to detect whether it causes layout conflicts, such as an overlong title bar or an unbalanced chart proportion.

[0069] If the test passes, the weights of the third fully connected layer of the local generator are updated, and a new version of the PPT is generated containing a bar chart comparing fast charging technologies. At the same time, the reduction in generation time is recorded, such as from an average of 45 seconds to 32 seconds.

[0070] The updated generator automatically triggers a feedback loop at the output end: when a user exports a PPT containing a fast-charging chart to PDF, the system captures this action as a positive feedback signal, generates a log file containing the weight adjustment amplitude, and returns it to the distillation channel; if the new parameters cause the PPT content generation failure rate to exceed 10% for five consecutive times, it rolls back to the previous stable version and marks the coefficient vector as a high-risk model.

[0071] The channel maintenance process initiates matrix optimization monthly: Analyzes historical transmission logs for basis matrices with compression error rates exceeding 15%, such as the sparse matrices generated by medical terminology transmission. Singular value reconstruction is used to control the error rate to less than 5%. The updated basis matrix is ​​then broadcast to all access units for the next transmission.

[0072] The cross-modal semantic distillation network uniformly processes text, voice, and visual data, fundamentally resolving the problem of information fragmentation. A temporal alignment attention mechanism is used to establish an emotional mapping between product videos and user reviews, improving the accuracy of multi-source information correlation. Combined with the dynamic optimization capabilities of the intent-generating adversarial network, autonomous iteration and updating of e-commerce copy are achieved. A dual-loop feedback engine automatically enhances core selling point features based on real-time user behavior and market sentiment, ensuring that generated content remains highly appealing. The speech analysis unit uses a dual verification mechanism of voiceprint and semantics. When a serious conflict between emotional expression and textual content is detected, contextual backtracking is used to accurately identify the true user intent, reducing the misjudgment rate. The design draft generation process is verified in stages based on structural specifications and aesthetic semantics, effectively balancing platform compliance and creative freedom.

[0073] The federated skill network architecture enables efficient knowledge transfer. The buzzword model of the e-commerce copywriting unit is converted into a creative vector for the design unit through low-rank compression, reducing the cost of integrating new functional modules. The asset-aware backup engine, based on a value-tiered strategy, performs differential neural coding compression on intelligently generated design drafts and documents, reducing storage overhead while preserving complete semantics. The blockchain verification module ensures the immutability of backup files through bit-level data comparison and a multi-version collaborative repair mechanism. The closed-loop feedback system of the knowledge distillation channel captures user behavior to drive parameter self-optimization, fostering the continuous evolution of system capabilities. Ultimately, a full-link closed loop covering data understanding, content generation, and asset protection is constructed, comprehensively improving the efficiency of digital content production.

[0074] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0075] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal AI agent and backup method based on deep learning, characterized in that: The steps include: Step S1: The multimodal cognitive center receives heterogeneous input data of text, voice, image and video; Step S2: The dynamic content generation system generates e-commerce copywriting, meeting minutes, design drawings, and presentations based on the output of the cognitive center; Step S3: The asset-aware backup engine performs encrypted and compressed backup of the content output by the generation system and the original input data; Step S4: the output end of the cognitive center is connected to the input end of the generation system, and the output end of the generation system is connected to the input end of the backup engine.

2. The multimodal AI agent and backup method based on deep learning according to claim 1, characterized in that: The dynamic content generation system includes an e-commerce copywriting generation unit, a speech analysis unit, a design generation unit, and a PPT generation unit; The e-commerce copywriting generation unit captures product details pages, user reviews, and video introductions, and dynamically updates product titles and promotional copywriting; The speech analysis unit transcribes sales calls and meeting recordings, and outputs customer demand analysis reports and structured meeting minutes; The design generation unit generates cross-platform visual materials according to natural language instructions; The PPT generation unit automatically generates training materials, marketing plans and project report documents based on the subject description.

3. The multimodal AI agent and backup method based on deep learning according to claim 1, characterized in that: The multimodal cognitive center in step S1 performs the following steps in order: Adaptive coding stage: extracting underlying features through text encoder, speech encoder and video encoder respectively; Cross-modal association stage: Use the temporal alignment attention mechanism to associate feature vectors of different modalities and establish a sentiment mapping relationship between product videos and user reviews; Content synthesis stage: The associated features are input into the intent generative adversarial network, the generator outputs the draft content, and the discriminator evaluates the content value based on the human cognitive model.

4. The multimodal AI agent and backup method based on deep learning according to claim 2, characterized in that: The e-commerce copywriting generation unit includes: Deploy a dual-loop feedback engine, with the inner loop monitoring the click-through rate of copywriting through real-time tracking data, and the outer loop combining market sentiment analysis results; When the copywriting effect is detected to be declining, the key selling point features in the user behavior data are strengthened based on the gradient reversal layer; The generator updates the copy content online based on the enhanced features.

5. The multimodal AI agent and backup method based on deep learning according to claim 2, characterized in that: The voice analysis unit performs two-factor authentication: Simultaneously extract the semantic features of the speech transcription text and the emotional features of the voiceprint of the original recording; When the difference between the voiceprint emotional feature and the text semantic confidence exceeds the preset threshold, the multi-channel recalibration module is triggered to perform secondary analysis; The output end is connected to the generation interface of the customer pain point analysis report and the meeting action item list.

6. The multimodal AI agent and backup method based on deep learning according to claim 2, characterized in that: The design generation unit adopts a phased assessment: Structural verification stage: Verify the component layout of the design draft through the preset typesetting rule engine; Aesthetic verification stage: Use the visual semantic matching model to calculate the consistency score of the design draft image and text; When both phases of evaluation are passed, the final design draft is sent to the backup engine.

7. The multimodal AI agent and backup method based on deep learning according to claim 1, characterized in that: The process of the asset-aware backup engine in step S3 performing encrypted compression backup on the content output by the generation system and the original input data includes: The file identification module marks the value hierarchy of original data, intermediate processed data and final output based on metadata; The core asset compression module uses a neural differential encoder to extract the semantic features of design drawings and PPT documents, achieving high-compression storage; The backup verification module is embedded in the blockchain light node, and bit-level data integrity comparison is performed through the hash chain during incremental backup.

8. The multimodal AI agent and backup method based on deep learning according to claim 7, characterized in that: The process of extracting semantic features of design drawings and PPT documents using a neural differential encoder includes: The encoding end extracts the deep semantic feature vector of the file; Compress the feature vector through the feature distillation layer; The decoding end reconstructs visual elements and text content based on semantic feature vectors.

9. The multimodal AI agent and backup method based on deep learning according to claim 1, characterized in that: The e-commerce copywriting generation unit, speech analysis unit, design generation unit, and PPT generation unit in the dynamic content generation system constitute a federated skill network; each unit shares the feature extraction layer of the multimodal cognitive hub, but has independent generator fine-tuning parameters; model parameters are exchanged between units through a knowledge distillation channel, where the hot word model of the e-commerce copywriting unit is converted into a creative prompt vector for the short video script unit.

10. The multimodal AI agent and backup method based on deep learning according to claim 9, characterized in that: The knowledge distillation channel includes: Low-rank matrix decomposition is used to compress the model parameters for transmission; The parameter receiving unit updates the local generator according to the new parameters; The updated generator output is fed back to the parameter adjustment interface of the multimodal cognitive center.

Citation Information

Patent Citations

  • Method for safe data backup, client side and cloud server side based on alliance chain

    CN106919476A

  • Behavior scheme determination method and device based on large model, electronic equipment and medium

    CN119228447A

  • Zero sample learning-based sonar image recognition method and system

    CN119851009A

  • Super employee artificial intelligence system

    CN120069824A

  • Multimedia information generation method and apparatus, and computer-readable storage medium

    WO2024061073A1