Intelligent material recommendation method and system based on image communication

By leveraging generative AI and image communication technologies, combined with multi-level feature extraction and multi-dimensional matching, the problem of disconnect between material recommendation and transmission has been solved, enabling personalized material generation and efficient transmission to meet the creative needs of multiple scenarios.

CN120856940APending Publication Date: 2025-10-28HANGZHOU YUNZHI CHUANGXIN NETWORK CO LTD

Patent Information

Application Number
CN202511340864.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing material recommendation technologies ignore the deep semantics of image content, disconnect between recommendation and transmission, have limited material libraries, and cannot meet the needs of personalized creation, leading users to manually adjust or abandon creation.

Method used

By integrating generative AI, combining image communication technology and intelligent recommendation logic, we can achieve multi-level feature extraction and multi-dimensional matching, generate materials that are highly consistent with user needs, and adopt adaptive compression and progressive transmission strategies to build a deep association between needs, images and materials.

Benefits of technology

It enables personalized material recommendations, improves creation efficiency and completeness, meets personalized needs in multiple scenarios, and optimizes transmission efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856940A_ABST
    Figure CN120856940A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent material recommendation method and system based on image communication, and belongs to the field of image communication and intelligent recommendation, and the system comprises a demand analysis and image collection module which is used for analyzing a user demand and collecting and transmitting a reference image; the image feature deep analysis module is used for extracting multi-level features of the reference image; the material library feature index module is used for material characterization and index establishment; the intelligent material generation and collaborative screening module is used for calling a generation model to generate candidate materials and screening an optimal generation material through a multi-dimensional scoring system; the intelligent matching recommendation module is used for performing multi-dimensional matching on the materials in the material library and the optimal generated materials to generate a recommendation list; according to the method, candidate materials can be dynamically generated based on reference image features and demand keywords through fusion of generative AI, and multi-scene personalized demands such as personal life records, enterprise commercial propaganda and mobile terminal short videos are covered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image communication and intelligent recommendation technology, and in particular to an intelligent material recommendation method and system based on image communication. Background Technology

[0002] In the process of content creation such as video editing, users' needs for materials (such as video clips, images, audio, special effects, etc.) are becoming increasingly diverse, but existing material recommendation technologies have obvious limitations: It is recommended to rely on text tag matching and ignore the deep semantics of image content (such as scene atmosphere, color style, and emotional tendency). The failure to incorporate the image features of real-time user editing resulted in recommended materials being out of sync with the style of existing content. The material transmission and recommendation processes are disconnected; image communication only focuses on transmission efficiency and is not linked to the recommendation logic. The recommended scope is limited to the existing material library. When there is no material in the material library that matches the user's personalized needs (such as "a shot of waves with specific lighting under the sunset"), the creative needs cannot be met, which causes the user to manually adjust or give up personalized creation, reducing editing efficiency.

[0003] Therefore, there is a need for an intelligent material recommendation method and system based on image communication, which integrates image communication technology, intelligent recommendation logic and generative AI to break through the boundaries of the material library and achieve a dual-track material supply of "accurate recommendation + dynamic generation". Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides an intelligent material recommendation method and system based on image communication. It solves the problems of disconnect between material recommendation and image content, separation of transmission and recommendation, and limited scope of material library in existing technologies. By integrating generative AI, it achieves a dual-track output of "intelligent generation + accurate recommendation", further improving the personalization and coverage of material recommendations.

[0005] Technical Solution: To solve the above-mentioned technical problems, according to one aspect of the present invention, more specifically, a method for intelligent material recommendation based on image communication, comprising the following steps: S1. Requirements Analysis and Image Acquisition Receive user editing request text (such as "Add a shot that matches the existing sunset seascape style and includes a surfer silhouette, paired with relaxing wave audio"), and extract request keywords (scene: sunset seascape, element: surfer silhouette, emotion: relaxed, material type: video shot, audio) through a natural language processing model (such as BERT).

[0006] Collect reference image data from the user's current editing project: Obtain 3-5 key frame sequences from the user's end through the improved RTP protocol. During transmission, an adaptive compression strategy is adopted (JPEG2000 compression ratio is 1:8 when the network bandwidth is ≥10Mbps, and the compression ratio is adjusted to 1:20 when the bandwidth is <5Mbps) to balance image quality and transmission efficiency.

[0007] S2, Image Feature Depth Analysis Preprocessing: The quality of the reference image is optimized using the BM3D noise reduction algorithm, and the image is enhanced to 1080P / 4K (matching the resolution of the user's editing project) using the EDSR super-resolution model to improve the accuracy of feature extraction.

[0008] Multi-level feature extraction: Basic features: HSL color histogram (75% warm tones), LBP texture (60% wave pattern), motion trajectory (wave undulation speed 0.5m / s); High-level semantic features: Identify scenes (beach sunset), objects (waves, sunset), and emotional tendencies (relaxed and peaceful) using the CLIP model; Style characteristics: Art style tags (natural realism) are extracted using the NeuralStyleTransfer model.

[0009] S3, Material Library Feature Index Material Characterization: For image / video materials in the material library, extract basic, semantic, and stylistic features consistent with the reference image to generate a 1024-dimensional feature vector; for audio materials, convert them into emotional feature vectors through Mel spectrum analysis (the high-frequency energy accounts for 35% of the lighthearted style) and associate them with the "beach" scene tag.

[0010] Dynamic index library: It uses Milvus vector database to store material feature vectors, and supports real-time updates (featureization and indexing of new materials are completed within 2 minutes) and millisecond-level retrieval.

[0011] S4, Intelligent Material Generation and Collaborative Screening Candidate material generation: The StableDiffusion diffusion model is called, and the multi-level features of the reference image extracted in step S2 (75% warm color tone, beach sunset scene, natural realistic style) and user demand keywords (surfer silhouette) are input to generate 3 sets of video shot candidate materials for "surfer silhouette at sunset" (1080P resolution, 30fps frame rate). At the same time, the generation model is called to generate 2 sets of audio candidate materials for "relaxed waves" (sampling rate 44.1kHz).

[0012] Multi-dimensional rating: Semantic consistency: The semantic similarity between the generated shots and the reference images is calculated using the BLIP-2 model (e.g., the similarity of the generated shots in the first group is 93%, the second group is 88%, and the third group is 82%). Style compatibility: Comparing the style feature vectors of the generated footage with those of the reference images, the style matching degree of Group 1 and Group 2 is 90%, while that of Group 3 is 80%. Format Compliance: All three generated shots conform to the user's 1080P / 30fps format, achieving a compliance score of 100%. Optimal selection: Calculate the comprehensive score based on "semantic consistency (40% weight) + style compatibility (40% weight) + format compliance (20% weight)". Group 1 (93%×0.4+90%×0.4+100%×0.2=91.2 points) and Group 2 (88%×0.4+90%×0.4+100%×0.2=89.2 points) are selected as the top two as the optimal generated materials. Similarly, select one group of optimal generated audio materials from the candidate audio materials.

[0013] S5, Intelligent Matching Recommendation Multi-dimensional matching: The "Sunset Seascape" video footage (Top-3) from the media library and the two sets of optimal generated shots from step S4 are included in the matching pool. The "Relaxing Waves" audio footage (Top-1) from the audio media library and one set of optimal generated audio are included in the matching pool. The following calculations are performed: Semantic matching: Cosine similarity between the keyword "surfer silhouette" and the semantic vector of the source material (92% similarity for generated shot 1, 85% similarity for shot 1 in the source material library); Visual matching: Weighted matching of warm color features of reference image with color features of material (user's historical preference "color priority", color weight increased to 50%, generated shot 1 color similarity 90%, material library shot 1 similarity 82%). Compatibility matching: Verify the copyright status (all are commercially usable) and format (1080P / 30fps) of all materials. Recommendation list generation: Generate Top-3 video clips (generated shot 1, clip 1 from the material library, and generated shot 2) and Top-2 audio clips (generated audio 1 and clip 1 from the material library) in descending order of overall matching degree, with the recommendation reasons (such as "generated shot 1: 93% semantic matching degree with the reference image sunset seascape, contains surfer silhouette elements, and the format is directly adapted").

[0014] S6. Recommendation Result Transmission and Interaction Optimized transmission: Video materials are transmitted in a progressive manner (a 720P preview version is transmitted first, completed within 10 seconds; a 1080P high-definition version is transmitted after user confirmation), and audio materials are encoded using ABR (bitrate of 192kbps when bandwidth ≥ 5Mbps, adjusted to 128kbps when bandwidth < 3Mbps).

[0015] User feedback: When users marked "Shot 1 generated has a perfect style" and "Shot 1 in the material library lacks elements", the system increased the semantic feature weight of the generated material by 15% and the "element matching" weight of the material library material by 10%, and iteratively optimized the matching model.

[0016] According to another aspect of the present invention, and more specifically, an intelligent material recommendation system based on image communication, comprising: Requirements Analysis and Image Acquisition Module: Includes a requirements text parsing unit (BERT model deployment) and a reference image acquisition and transmission unit (improved RTP protocol implementation); Image feature depth analysis module: includes image preprocessing unit (BM3D+EDSR) and multi-level feature extraction unit (CLIP+NeuralStyleTransfer); Material library feature index module: includes material preprocessing unit (feature algorithm) and vector indexing unit (Milvus database); Intelligent material generation and collaborative screening module: includes generation model unit (StableDiffusion + audio generation model), multi-dimensional scoring unit (BLIP-2 + style comparison + format verification), and conditional constraint unit (format / emotional constraint input). The intelligent matching and recommendation module includes a multi-dimensional matching unit (semantic + visual + compatibility) and a recommendation list generation unit (ranked by comprehensive matching degree). Recommendation result transmission and interaction module: includes optimized transmission unit (progressive transmission + ABR) and user feedback processing unit (weighted iteration).

[0017] The beneficial effects of the intelligent material recommendation method and system based on image communication of the present invention are as follows: (1) This invention integrates generative AI to dynamically generate candidate materials based on reference image features and demand keywords, covering personalized needs in multiple scenarios such as "personal life records, corporate commercial promotion, and mobile short videos". For example, in the corporate promotional video scenario, it can accurately generate business meeting room shots containing the brand's deep blue VI color; in the food short video scenario, it can generate rare close-up shots of tomato cross sections, completely breaking through the boundaries of the material library, allowing users to achieve personalized creation without manual adjustment, and greatly improving the completeness and uniqueness of creation.

[0018] (2) This invention constructs a deep association between "demand-image-material" through multi-level feature extraction (basic features + high-level semantic features + style features) and a multi-dimensional matching mechanism. In the feature extraction stage, the CLIP visual large model is used to identify scenes, objects and emotions, and the NeuralStyleTransfer model is used to extract artistic style tags; in the matching stage, semantic matching (cosine similarity calculation), visual matching (weighted feature comparison), and compatibility matching (format and copyright verification) are integrated to ensure that the recommended materials are highly consistent with the reference images in terms of semantics, style and emotion.

[0019] (3) This invention achieves deep integration of communication and recommendation through an improved communication protocol and dynamic transmission strategy: In the acquisition stage, the JPEG2000 compression ratio is dynamically adjusted based on network bandwidth to balance image quality and transmission efficiency; in the recommendation result transmission stage, progressive transmission and adaptive bit rate encoding are adopted to adapt to different devices and bandwidth environments. For example, in a 4G mobile environment, 720P video material is first transmitted in 480P preview version, and then in high-definition version after confirmation, to avoid transmission lag caused by insufficient bandwidth; in enterprise professional scenarios, the compression ratio of 4K material is set to 1:8 to retain details, and the transmission time is only 12 seconds. At the same time, the matching weight is iteratively optimized through user feedback to continuously optimize the recommendation effect, forming a closed loop of "transmission-recommendation-feedback", which significantly improves the user experience. Attached Figure Description

[0020] The present invention will now be described in further detail with reference to the accompanying drawings and specific implementation methods.

[0021] Figure 1 This is a system structure diagram of the present invention; Figure 2 Flow chart of the method of the present invention. Detailed Implementation

[0022] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.

[0023] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] Example 1 Reference Figure 1 A method and system for intelligent content recommendation based on image communication, targeting the scenario of a user editing a "summer travel vlog (highlighting 'personalized sunset and sea view')," includes the following steps: Requirements Analysis and Image Acquisition: The user inputs "I need a sunset seascape shot with coconut tree shadows, paired with relaxing audio featuring birdsong". The keywords are analyzed (scene: sunset seascape, element: coconut tree shadows, emotion: relaxing, audio element: birdsong). Two edited sunset seascape keyframes are acquired. With a bandwidth of 4Mbps, the JPEG2000 compression ratio is set to 1:15, and the transmission time is 8 seconds.

[0025] Image feature in-depth analysis: After BM3D denoising, EDSR super-resolution up to 1080P; basic features (warm tone ratio 80%, coconut tree shadow texture ratio 30%), semantic features (beach sunset, coconut tree, relaxed), and style features (natural realistic style) are extracted.

[0026] Feature index of the material library: There are 50 video materials in the material library with the theme of "sunset sea view". After feature extraction, the material with the highest semantic matching degree with the reference image is 86% (without coconut tree shadow elements); There are 20 audio materials in the "relaxing birdsong" category. The highest emotional matching degree is 88%.

[0027] Intelligent material generation and collaborative selection: Generate 3 sets of "sunset seascape with coconut tree shadows" video footage (1080P / 30fps) and 2 sets of "relaxing audio with birdsong"; Scoring: Generated Shot 1 (Semantic similarity 94%, style matching 92%, format compliance 100%, overall score 92.8), Generated Shot 2 (90% / 91% / 100%, score 90.4), Generated Audio 1 (Emotional matching 95%, format compliance 100%, score 97.5); Select the top-2 videos and top-1 audio as the best source material.

[0028] Intelligent matching recommendation: The matching pool contains the top-3 videos (without coconut tree shadows) from the material library and the top-2 generated videos, the top-2 audios from the material library and the top-1 generated audios; the overall matching score is ranked as "generated shot 1 (92.8 points) > generated shot 2 (90.4 points) > material library shot 1 (86 points)" and "generated audio 1 (97.5 points) > material library audio 1 (88 points)", generating a recommendation list.

[0029] Transmission and Interaction: The 720P generated shot 1 preview version was transmitted in 5 seconds, and the 1080P HD version was transmitted in 20 seconds after user confirmation; user feedback that "generated shot 1 elements are accurate" and "generated audio 1 bird song volume is moderate" led the system to increase the "element matching" weight of the generated materials by 20%, and generated materials will be given priority in future recommendations.

[0030] This embodiment is aimed at a lightweight creation scenario for individual users. The core requirements are "emotional expression + personalized visual elements + medium bandwidth adaptation". Typical devices are home computers / mid-range mobile phones. The creative goal is to produce life record content with a consistent style and emotional relevance. Focusing on "personal emotional needs," the technical solution does not require commercial attributes such as brand VI adaptation or professional copyright verification. It emphasizes the synergy of "emotion, vision, and audio" rather than extreme precision or mobile adaptation.

[0031] Example 2 Reference Figure 1 A method and system for intelligent material recommendation based on image communication, targeting the editing scenario of corporate promotional videos (technology product launches), includes the following steps: Requirements Analysis and Image Acquisition User input: "Add a business meeting room shot that matches the style of existing smart terminal product displays (must include the brand's deep blue VI color elements), paired with low-saturation, tech-inspired background music, adapted for 4K / 60fps professional editing projects." The requirement analysis unit extracts keywords using the BERT model: scene (business meeting room), style (tech minimalist style), visual elements (brand's deep blue VI color), material type (video shot, background music), and format (4K / 60fps).

[0032] Reference Image Acquisition: Three keyframes of a "close-up of a smart terminal product" (including the brand's dark blue logo and metallic body) edited by the user were acquired and transmitted via an improved RTP protocol. The user's terminal was detected as a professional workstation (network bandwidth 15Mbps), and the JPEG2000 compression ratio was set to 1:8 to ensure that the metallic texture details were preserved after transmission. The transmission took 12 seconds.

[0033] Image Feature Deep Analysis Preprocessing: The BM3D noise reduction algorithm is used to eliminate reflective noise in the product close-up, and the image is enhanced to 4K (matching the editing project format) using the EDSR super-resolution model to avoid resolution banding in the subsequent generated material.

[0034] Multi-level feature extraction: Basic features: HSL color histogram (brand deep blue accounts for 45%, metallic silver gray accounts for 30%), LBP texture (brushed texture of product body accounts for 60%), motion trajectory (no dynamic, marked as "static product close-up"); High-level semantic features: Identify scenes (product display), objects (smart terminals), and emotional tendencies (professionalism, technological feel) through CLIP model. Style characteristics: Tags extracted using the NeuralStyleTransfer model (minimalist business style, cool-toned tech vibe).

[0035] Material Library Feature Index Material library search: There are 80 video materials in the "Business Meeting Room" category in the material library. After feature extraction, only 75% of the materials have the highest style matching degree with the reference images (no brand dark blue VI color, and mostly warm color tone); there are 30 "Technological Background Sounds". Through Mel spectrum analysis, the highest emotional matching degree is 82% (the high frequency energy ratio is too high, which can easily mask the product explanation sound).

[0036] Index Update: Prioritize the feature vectors of materials tagged "high-tech" in the material library to ensure they are used first during subsequent matching.

[0037] Intelligent material generation and collaborative screening Candidate material generation: The StableDiffusion diffusion model is called, and the multi-level features of the reference image are input (brand deep blue 45%, minimalist business style, 4K / 60fps format constraint) and the required keywords (business conference room, brand VI color) to generate 4 sets of video shots of "business conference room containing brand deep blue elements" (such as "deep blue curtains + silver gray conference table" scene); At the same time, the audio generation model (such as MusicGen) is called, and the constraints of "low saturation technology feel + suitable for product explanation" are input to generate 3 sets of background sound (low frequency electronic sound effects account for 60% to avoid interfering with human voices).

[0038] Multi-dimensional scoring (weighted by "brand suitability 40% + style consistency 30% + format compliance 30%)): Video Candidate 1: Brand Deep Blue accounts for 42% (adaptability 93%), style matching 90%, format compliance 100%, overall score 91.2; Video Candidate 2: Brand Deep Blue accounts for 38% (adaptability 84%), style matching degree 92%, format compliance 100%, overall score 88.4; Audio Candidate 1: 95% matching of low saturation and technological feel, 90% compatibility with product explanation audio, 100% compliance with format (48kHz sampling rate), overall score 93.5; Optimal selection: Select the top-2 videos (candidate 1, candidate 2) and the top-1 audio (candidate 1) as the optimal generated materials to ensure that the brand style and format are fully compatible.

[0039] Intelligent matching recommendation Matching pool construction: The top-3 videos (75%-80% matching degree) of "Business Meeting Room" in the material library and the generated top-2 videos, and the top-2 videos (82%-85% matching degree) of "Technological Background Sound" in the material library and the generated top-1 audio are included in the matching pool.

[0040] Multi-dimensional matching: Semantic matching: The semantic similarity between the requirement "brand VI color" and the generated video 1 is 93%, while that between the video 1 in the material library and only 65%. Visual matching: Based on the user's historical preference for "brand color priority", the color weight was set to 50%, resulting in a 92% color similarity for the generated video 1 and a 70% similarity for the video 1 in the material library; Compatibility: All materials have passed copyright verification (commercial license), and the formats are 4K / 60fps (video) and 48kHz (audio). Recommended list: Sorted by overall matching degree as "Generated Video 1 (91.2 points) > Generated Video 2 (88.4 points) > Material Library Video 1 (80 points)" and "Generated Audio 1 (93.5 points) > Material Library Audio 1 (85 points)", with the recommendation reasons (such as "Generated Video 1: contains the brand's dark blue VI color elements, matches the product display style 90%, and is directly adapted to the project in 4K / 60fps").

[0041] Transmission and Interaction Optimized transmission: Video is transmitted progressively (a 1080P preview version is transmitted first, which takes 15 seconds; a 4K high-definition version is transmitted after user confirmation, which takes 40 seconds); audio uses ABR encoding (bitrate of 256kbps when bandwidth is stable, ensuring low-saturation sound detail).

[0042] User feedback: Users marked "Generated video 1 with accurate brand elements" and "Generated audio 1 without interfering with the narration". The system will increase the weight of "brand adaptability" of the generated materials by 25%, and will prioritize strengthening the matching of brand VI colors in subsequent corporate promotional video editing.

[0043] This example is geared towards commercial creative scenarios for enterprises / professional editors. The core requirements are "brand VI consistency + professional format adaptation (4K / 60fps) + copyright compliance". The typical equipment is a professional workstation. The creative goal is to produce business content that conforms to the brand image and is suitable for large-screen playback at press conferences. Focusing on "professional commercial attributes," this solution needs to balance brand consistency, high format accuracy, and copyright compliance. It strengthens "brand characteristic constraints" and "professional parameter adaptation," which distinguishes it from the emotional emphasis in personal scenarios and the lightweight requirements of short video scenarios.

[0044] Example 3 Reference Figure 1 A method and system for intelligent content recommendation based on image communication, targeting the editing scenario of food tutorials (Chinese home-style dish - scrambled eggs with tomatoes) on short video platforms, includes the following steps: Requirements Analysis and Image Acquisition User input: "Add close-up shots of the cross-section of diced tomatoes and bubbly eggs being scrambled, accompanied by crisp chopping sounds and sizzling sounds, adapted to 720P / 30fps short video format." The requirement analysis unit extracted the following keywords: scene (kitchen cooking), elements (close-up shots of diced tomatoes and scrambled eggs), sound effects (chopping sounds and sizzling sounds), and format (720P / 30fps).

[0045] Reference Image Acquisition: Two keyframes of "tomatoes being washed and ready to be cut" (containing bright red tomatoes and a white porcelain plate) already captured by the user were acquired and transmitted from the mobile device via an improved RTP protocol. The user's network bandwidth was detected to be 3Mbps (mobile 4G environment). The JPEG2000 compression ratio was set to 1:18 to ensure transmission was completed within 8 seconds, while preserving the tomato's color details.

[0046] Image Feature Deep Analysis Preprocessing: Lightweight BM3D noise reduction algorithm (adapted to mobile computing power) is used to eliminate image reflections, and tiny-EDSR super-resolution model (parameter compression 60%) is used to enhance the image to 720P to avoid blurry generated material.

[0047] Multi-level feature extraction: Basic features: HSL color histogram (65% tomato red, 25% porcelain white), LBP texture (70% tomato skin texture), motion trajectory (static, marked as "ingredient preparation"); High-level semantic features: Identify scenes (kitchen), objects (tomatoes, porcelain plates), and sentiment (lifelike, clear) through the lightweight CLIP model (MobileCLIP); Style characteristics: Tags were extracted by simplifying the NeuralStyleTransfer model (realistic kitchen style, high color saturation).

[0048] Material Library Feature Index Material library search: There are 120 videos related to "scrambled eggs with tomatoes" in the material library, only 20 of which are close-ups of "cutting tomatoes into chunks", and there are no "cross-section close-ups" (most are side shots), with a highest matching degree of 70%; there are 50 videos related to "cooking sound effects", only 8 of which are "crisp sound of cutting tomatoes", with a matching degree of 80% (high frequency bands are not prominent enough).

[0049] Index optimization: Set the tags "Ingredient close-up" and "Cooking sound effects" as priority search terms on mobile devices to shorten search time.

[0050] Intelligent material generation and collaborative screening Candidate material generation: The lightweight diffusion model (StableDiffusionXL-Turbo, which improves generation speed by 50%) is called. The input reference image features (65% tomato red, realistic style, 720P / 30fps) and the required keywords (cross-section of tomato chunks, bubbling eggs when stir-fried) are used to generate 3 sets of close-up shots (5 seconds each, adapted to the requirements of short video clips); The audio generation model (AudioGen-Lite) is called. The input constraint "high-frequency crisp sound of cutting tomatoes + low-frequency sizzling sound of stir-frying" is used to generate 2 sets of cooking sound effects (3 seconds each).

[0051] Multi-dimensional scoring (weighted by "detail matching degree 40% + mobile adaptation 30% + duration compliance 30%)): Shot Candidate 1 (Tomato Cross-section): Detail matching 95% (tomato seeds are visible), compatibility 100% (720P / 30fps), duration compliance 100%, overall score 98; Shot Candidate 2 (Scrambled Eggs): Detail matching 92% (visible bubble texture), adaptability 100%, duration compliance 100%, overall score 94.8; Sound effect candidate 1: The sound of chopping vegetables has a high frequency ratio of 45% (crisp), the sound of stir-frying has a low frequency ratio of 55% (mellow), and the compatibility is 100% (128kbps bit rate), with an overall score of 96. Optimal selection: Select the top-2 shots (candidate 1, candidate 2) and the top-1 sound effect (candidate 1) as the optimal generated materials to ensure that the details are adapted to the needs of mobile devices.

[0052] Intelligent matching recommendation Matching pool construction: The top-2 shots (70%-75% matching degree) of "Tomato Slices" in the material library and the generated top-2 shots, and the top-2 shots (80%-82% matching degree) of "Cooking Sound Effects" in the material library and the generated top-1 sound effects are included in the matching pool.

[0053] Multi-dimensional matching: Semantic matching: The requirement "cross-sectional close-up" has a 95% similarity to the generated shot 1, while shot 1 in the material library has only a 60% similarity. Visual matching: Mobile users prefer "vibrant colors", so the color weight is set to 40%. The generated shot 1 has a tomato red similarity of 93%, while the similarity of shot 1 in the material library is 80%. Compatibility: All media formats are 720P / 30fps (video) and 128kbps (audio), compatible with mobile editing software; Recommended list: sorted by overall matching degree as "Generated Shot 1 (98 points) > Generated Shot 2 (94.8 points) > Footage Library Shot 1 (75 points)" and "Generated Sound Effect 1 (96 points) > Footage Library Sound Effect 1 (82 points)", with the recommendation reasons (such as "Generated Shot 1: The close-up of the tomato cross section is clear and the color is consistent with the existing tomato image, so it can be directly inserted into the tutorial steps").

[0054] Transmission and Interaction Optimized transmission: Video uses a progressive transmission of "ultra-light preview version (480P, completed in 3 seconds) + high-definition version (720P, completed in 10 seconds)"; audio uses ABR encoding (automatically reduced to 96kbps when bandwidth is <2Mbps to ensure no stuttering).

[0055] User feedback: Users marked "Generated Shot 1 close-up is practical" and "Generated Sound Effect 1 is immersive". The system will increase the weight of "Ingredient detail matching" by 20%, and will recommend generating close-up materials first in future food tutorials.

[0056] For mobile creation scenarios targeting the general public, the core requirements are "detailed close-ups of food ingredients + low bandwidth adaptation + short-duration footage (3-5 seconds)", with typical devices being mid-to-low-end mobile phones (4G environment), and the creative goal is to produce intuitive and easy-to-insert short video tutorial clips. Focusing on "mass mobile needs", the technical solution is based on "lightweight, low bandwidth, and short duration", which is different from the high precision of enterprise scenarios and the medium bandwidth adaptation of personal scenarios. It is specifically designed to solve the pain points of limited computing power / bandwidth and lack of detailed close-up shots on mobile devices.

[0057] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for intelligent material recommendation based on image communication, characterized in that, Includes the following steps: S1. Requirements Analysis and Image Acquisition: Analyze the user's editing requirements text, acquire the reference images provided by the user through the image communication protocol, and transmit them. S2. Deep analysis of image features: Preprocess the reference image to extract basic features, high-level semantic features and style features; S3. Feature Index of Material Library: Perform feature processing on the materials in the material library and establish a feature vector index library; S4. Intelligent material generation and collaborative screening: Based on the multi-level features of the reference image extracted in step S2, the generation model is called to generate candidate materials, and the optimal generated materials are screened through a multi-dimensional scoring system; S5. Intelligent matching and recommendation: Combining demand features and reference image features, the materials in the material library in step S3 are matched with the optimal generated materials in step S4 in multiple dimensions to generate a recommendation list; S6. Recommendation Result Transmission and Interaction: Transmit recommendation materials through an optimized image communication strategy, receive user feedback, and iteratively optimize.

2. The intelligent material recommendation method based on image communication according to claim 1, characterized in that, In step S1, the demand analysis uses a natural language processing model to extract keywords, the image acquisition and transmission uses an improved RTP protocol, and the JPEG2000 compression ratio is dynamically adjusted according to the network bandwidth.

3. The intelligent material recommendation method based on image communication according to claim 1, characterized in that, The image features in step S2 include: Basic features: color histogram, texture features, motion trajectory; High-level semantic features: scenes, objects, and emotional tendencies identified through large visual models; Style features: Art style tags extracted through a style transfer model.

4. The intelligent material recommendation method based on image communication according to claim 1, characterized in that, The intelligent material generation and collaborative screening in step S4 specifically includes: Model generation: Using a diffusion model, the multi-level features of the reference image extracted in step S2 are input to generate candidate materials that are styled and semantically compatible with the reference image. The candidate materials include images, video clips, and audio. A multi-dimensional scoring system is constructed: candidate generated materials are scored from three dimensions: semantic consistency, style compatibility, and format compliance. Semantic consistency: Calculate the semantic similarity between the generated material and the reference image using a visual-language model; Style compatibility: Compare the style feature vectors of the generated material with those of the reference image; Format compliance: Verify the compatibility of the generated footage with the user's existing editing projects in terms of resolution and frame rate; Optimal generated material selection: Select candidate materials with a score ≥ preset threshold. If there are multiple materials that meet the criteria, select the Top-1 to Top-2 materials in descending order of score.

5. The intelligent material recommendation method based on image communication according to claim 4, characterized in that, When generating audio material, the emotional features of the reference image are converted into audio spectrum constraints through a cross-modal model to ensure that the emotional tendency of the generated audio is consistent with that of the reference image.

6. The intelligent material recommendation method based on image communication according to claim 1, characterized in that, The intelligent matching recommendation in step S5 includes semantic matching, visual matching, and compatibility matching, wherein: The semantic matching is achieved based on a similarity algorithm between the demand keywords and the semantic features of the materials; The visual matching is achieved based on a weighted calculation of multi-dimensional visual features of the reference image and the source material. The compatibility matching includes format adaptation verification and copyright legality verification; During matching, materials from the material library and the best generated materials are included in the matching pool, and a recommendation list is generated in descending order of overall matching degree.

7. An intelligent material recommendation system based on image communication, characterized in that, include: Requirement Analysis and Image Acquisition Module: Used to analyze user requirements and acquire and transmit reference images; Image feature deep analysis module: used to extract multi-level features from the reference image; Material Library Feature Index Module: Used for material feature generation and index creation; Intelligent material generation and collaborative screening module: used to call the generation model to generate candidate materials, and to screen the best generated materials through a multi-dimensional scoring system; The intelligent matching and recommendation module is used to perform multi-dimensional matching between materials in the material library and the best generated materials to generate a recommendation list; Recommendation Result Transmission and Interaction Module: Used to optimize the transmission of recommendation materials and process user feedback.

8. The intelligent material recommendation system based on image communication according to claim 7, characterized in that, The intelligent material generation and collaborative screening module includes: Generative Model Unit: Deploys a diffusion model to support the generation of source material based on reference image features; Multi-dimensional scoring unit: Integrates BLIP-2 model, style feature comparison algorithm and format verification logic to achieve semantic, style and format scoring of generated materials; Conditional constraint unit: The format parameters of the user's editing project and the emotional features of the reference image are used as input constraints for the generation model to ensure that the generated materials can be used without secondary processing.

9. The intelligent material recommendation system based on image communication according to claim 7, characterized in that, The image feature deep analysis module includes an image preprocessing unit and a multi-level feature extraction unit; The recommendation result transmission and interaction module adopts progressive transmission and adaptive bit rate encoding technology.

Citation Information

Patent Citations

  • Image material recommendation method and device

    CN106355429A

  • Household appliance recommendation method and device, electronic equipment and storage medium

    CN118886984A

  • Image material generation method and device, medium and computing equipment

    CN119850789A

  • Text-to-image generation model evaluation method and system based on multi-modal large model

    CN120071055A

  • Character image generation method and device, electronic equipment and storage medium

    CN120374767A

Cited By

  • Multimedia material intelligent retrieval method and system based on digital multimedia

    CN121478994A

  • Intelligent Retrieval Method and System for Multimedia Materials Based on Digital Multimedia

    CN121478994B