Intelligent service system and method for customized cultural and creative products

Through deep learning-based image processing technology and cross-modal semantic interaction, the problem of understanding deviation and long cycle in cultural and creative product customization is solved, and customized effect image generation is achieved that quickly responds to personalized needs, improving user experience.

CN119784471BActive Publication Date: 2025-08-15BEIJING LIDINGDANG CREATIVE CULTURE CO LTD

Patent Information

Application Number
CN202411853056.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-08-15
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The communication and understanding deviation between users and designers in the traditional cultural and creative product customization model leads to the inconsistent product expectations, the design cycle is long and difficult to iterate quickly, and it is unable to efficiently respond to users' immediate needs.

Method used

Image processing technology based on deep learning is used to extract images from the product prototype library, and fine-grained cross-modal semantic interaction is carried out in combination with user custom requirements to generate customized effect images, including fine-grained semantic coding feature extraction of product prototype images, customized requirements semantic coding and cross-modal interaction optimization modulation.

Benefits of technology

Quickly respond to users' personalized customization needs, shorten the design cycle, improve customization efficiency, reduce costs, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784471B_ABST
    Figure CN119784471B_ABST
Patent Text Reader

Abstract

The present application discloses an intelligent service system and method for customization of cultural and creative products, which relates to the field of intelligent customization. The system extracts product prototype images from a product prototype library as initial materials, and uses deep learning-based image processing technology to perform semantic analysis on the prototype images of cultural and creative products, thereby capturing key design elements in the images. At the same time, combined with the customization requirements of cultural and creative products input by users, fine-grained cross-modal semantic interaction is performed between the product customization requirements and the product prototype images, thereby realizing customized feature modulation of the product prototype images according to the product customization requirements, and intelligently generating customized effect images of cultural and creative products based on this. This system can quickly respond to users' personalized customization requirements, shorten the design cycle, improve customization efficiency, reduce customization costs, and thus effectively enhance users' customization experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent customization, and more specifically, to an intelligent service system and method for customizing cultural and creative products. Background Art

[0002] With the rapid development of the cultural and creative industries, market demand for cultural and creative products, as important vehicles for cultural expression and dissemination, is growing. Consumers are also increasingly demanding personalized customization of these products, no longer content with cookie-cutter goods but preferring products that reflect their personal style and unique taste. This trend is driving the expansion of the market for customized cultural and creative product services and technological innovation.

[0003] In the field of customized cultural and creative products, the traditional customization model typically requires multiple, in-depth conversations between users and designers, with the designers making adjustments based on the user's explanations and needs. However, due to a lack of efficient and practical tools and technical support, this interaction between users and designers can lead to misunderstandings, resulting in the final product not meeting user expectations. Furthermore, this customization model has a relatively long design cycle, making rapid iteration difficult and ineffective in responding to users' immediate needs.

[0004] Therefore, an intelligent service system and method for customizing cultural and creative products is needed to solve the above technical problems. Summary of the Invention

[0005] In order to solve the above technical problems, this application is proposed.

[0006] According to one aspect of the present application, a method for intelligent service of customizing cultural and creative products is provided, which includes:

[0007] Extracting product prototype images from a product prototype library and simultaneously obtaining cultural and creative product customization requirements input by users, wherein the cultural and creative product customization requirements include aesthetic preferences, cultural background, and descriptions of specific occasion requirements;

[0008] Performing image semantic feature extraction on the product prototype image to obtain a fine-grained semantic encoding feature map of the product prototype image;

[0009] Standardizing and semantically encoding the cultural and creative product customization requirements to obtain a semantic encoding feature vector of the cultural and creative product customization requirements;

[0010] Based on the cultural and creative product customization demand semantic coding feature vector, fine-grained semantic interaction optimization modulation is performed on the product prototype image fine-grained semantic coding feature map to obtain a customization demand modulation product image feature inversion feature map;

[0011] Based on the customization requirements, the product image feature inversion feature map is modulated to generate a customization effect to obtain a customized effect image of the cultural and creative product.

[0012] Specifically, performing image semantic feature extraction on the product prototype image to obtain a fine-grained semantic encoding feature map of the product prototype image includes:

[0013] The product prototype image is input into a product prototype image feature extractor based on the ViT model to obtain a fine-grained semantic encoding feature map of the product prototype image.

[0014] Specifically, the cultural and creative product customization requirements are standardizedly expressed and semantically encoded to obtain a semantic encoding feature vector of the cultural and creative product customization requirements, including:

[0015] Inputting the cultural and creative product customization requirements into a customization requirement specification expression module based on a large language model to obtain a cultural and creative product customization requirement specification expression;

[0016] The cultural and creative product customization requirement specification expression is semantically encoded to obtain the cultural and creative product customization requirement semantic encoding feature vector.

[0017] Specifically, based on the cultural and creative product customization demand semantic coding feature vector, fine-grained semantic interaction optimization modulation is performed on the product prototype image fine-grained semantic coding feature map to obtain a customization demand modulation product image feature inversion feature map, including:

[0018] Performing autocorrelation coding on the semantic coding feature vector of the cultural and creative product customization demand to obtain an autocorrelation coding matrix of the cultural and creative product customization demand;

[0019] Performing local fine-grained semantic interaction coding on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices;

[0020] Based on the set of semantic coding matrices of the customization demand-product prototype image cross-modal prompt information, the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map are guided to perform cross-modal interactive optimization to obtain the customization demand modulated product image feature inversion feature map.

[0021] Specifically, the cultural and creative product customization demand semantic coding feature vector is autocorrelatedly coded to obtain a cultural and creative product customization demand autocorrelation coding matrix, including:

[0022] The product between the semantic coding feature vector of the cultural and creative product customization demand and the transposed vector of the semantic coding feature vector of the cultural and creative product customization demand is calculated to obtain the autocorrelation coding matrix of the cultural and creative product customization demand.

[0023] Specifically, local fine-grained semantic interaction coding is performed on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices, including:

[0024] Performing feature decoupling on the fine-grained semantic encoding feature map of the product prototype image to obtain a set of local semantic feature matrices of the product prototype image;

[0025] Performing a linear transformation on the cultural and creative product customization demand autocorrelation coding matrix to obtain a cultural and creative product customization demand query coding matrix and a cultural and creative product customization demand value coding matrix;

[0026] The local semantic feature matrices of each product prototype image in the set of the cultural and creative product customization demand query coding matrix, the cultural and creative product customization demand value coding matrix and the product prototype image local semantic feature matrix are respectively input into the cross-modal prompt information encoder based on the converter structure to obtain the set of the customization demand-product prototype image cross-modal prompt information semantic coding matrices.

[0027] Specifically, based on the set of the semantic coding matrix of the customization demand-product prototype image cross-modal prompt information, guiding the cultural and creative product customization demand semantic coding feature vector and the product prototype image fine-grained semantic coding feature map to perform cross-modal interactive optimization to obtain the customization demand modulated product image feature inversion feature map, including:

[0028] Inputting the set of the customization demand-product prototype image cross-modal prompt information semantic encoding matrices into the decoder-based information gating unit to obtain a set of customization demand-product prototype image cross-modal semantic interaction attention weights;

[0029] Inputting the set of the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image local semantic feature matrix into a cross-modal interaction unit to obtain a set of customization demand-product prototype image cross-modal interaction local feature matrices;

[0030] The set of attention weights of the cross-modal semantic interaction between the customization requirement and the product prototype image and the set of local feature matrices of the cross-modal interaction between the customization requirement and the product prototype image are input into the cross-modal interaction optimization unit to obtain the feature inversion feature map of the customization requirement modulated product image.

[0031] Specifically, based on the customization requirement, the product image feature inversion feature map is modulated to generate a customization effect to obtain a customized effect image of the cultural and creative product, including:

[0032] The customized demand modulated product image feature inversion feature map is input into a customized effect generator based on a diffusion model to obtain the customized effect image of the cultural and creative product.

[0033] According to another aspect of the present application, there is provided an intelligent service system for customizing cultural and creative products, comprising:

[0034] A product image customization demand acquisition module is used to extract product prototype images from the product prototype library and simultaneously acquire cultural and creative product customization demands input by users, wherein the cultural and creative product customization demands include aesthetic preferences, cultural background, and descriptions of demands for specific occasions;

[0035] A product prototype image semantic feature extraction module, configured to extract image semantic features from the product prototype image to obtain a fine-grained semantic encoding feature map of the product prototype image;

[0036] A customization requirement semantic coding module, configured to perform standardized expression and semantic coding on the customization requirement of the cultural and creative product to obtain a semantic coding feature vector of the customization requirement of the cultural and creative product;

[0037] A customization demand modulation module is used to perform fine-grained semantic interaction optimization modulation on the fine-grained semantic coding feature map of the product prototype image based on the semantic coding feature vector of the customization demand of the cultural and creative product to obtain a customization demand modulation product image feature inversion feature map;

[0038] The customized effect generation module is used to modulate the product image feature inversion feature map based on the customized requirements to generate customized effects to obtain customized effect images of cultural and creative products.

[0039] This application has at least the following technical effects:

[0040] Compared with the existing technology, the present application provides an intelligent service system and method for customization of cultural and creative products. It extracts product prototype images from a product prototype library as initial materials, and uses deep learning-based image processing technology to perform semantic analysis on the prototype images of cultural and creative products to capture the key design elements in the images. At the same time, combined with the cultural and creative product customization requirements input by the user, through fine-grained cross-modal semantic interaction between product customization requirements and product prototype images, customized feature modulation of product prototype images is achieved according to product customization requirements, and customized effect images of cultural and creative products are intelligently generated based on this. It can quickly respond to users' personalized customization needs, shorten the design cycle, improve customization efficiency, reduce customization costs, and thus effectively enhance the user's customization experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0042] Figure 1 This is a flowchart of an intelligent service method for customizing cultural and creative products according to an embodiment of the present application.

[0043] Figure 2 This is a flowchart for standardizing and semantically encoding the cultural and creative product customization requirements in the intelligent service method for cultural and creative product customization according to an embodiment of the present application to obtain a semantic encoding feature vector of the cultural and creative product customization requirements.

[0044] Figure 3 This is a flowchart of performing fine-grained semantic interactive optimization modulation on the fine-grained semantic coding feature map of the product prototype image based on the semantic coding feature vector of the cultural and creative product customization demand in the intelligent service method for customization of cultural and creative products according to an embodiment of the present application to obtain a customization demand modulated product image feature inversion feature map.

[0045] Figure 4 This is a block diagram of an intelligent service system for customizing cultural and creative products according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. While the drawings illustrate certain embodiments of the present disclosure, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0047] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0048] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0049] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0050] Specifically, in this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0051] With the rapid development of the cultural and creative industries, market demand for cultural and creative products, as important vehicles for cultural expression and dissemination, is growing. Consumers are also increasingly demanding personalized customization of these products, no longer content with cookie-cutter goods but preferring products that reflect their personal style and unique taste. This trend is driving the expansion of the market for customized cultural and creative product services and technological innovation.

[0052] In the field of customized cultural and creative products, the traditional customization model typically requires multiple, in-depth conversations between users and designers, with the designers making adjustments based on the user's explanations and needs. However, due to a lack of efficient and practical tools and technical support, this interaction between users and designers can lead to misunderstandings, resulting in the final product not meeting user expectations. Furthermore, this customization model has a relatively long design cycle, making rapid iteration difficult and ineffective in responding to users' immediate needs.

[0053] Therefore, in response to the above technical problems, the technical concept of this application is: extract product prototype images from the product prototype library as initial materials, and use deep learning-based image processing technology to perform semantic analysis on the prototype images of cultural and creative products to capture the key design elements in the images. At the same time, combined with the cultural and creative product customization needs input by the user, through fine-grained cross-modal semantic interaction between product customization needs and product prototype images, customized feature modulation of product prototype images is achieved according to product customization needs, and customized effect images of cultural and creative products are intelligently generated based on this, which can quickly respond to users' personalized customization needs, shorten the design cycle, improve customization efficiency, reduce customization costs, and thus effectively enhance the user's customization experience.

[0054] Figure 1 Flowchart of the intelligent service method for customizing cultural and creative products according to the embodiment of the present application. Figure 1As shown, the intelligent service method for customization of cultural and creative products according to the embodiment of the present application includes: S110, extracting product prototype images from a product prototype library, and at the same time obtaining cultural and creative product customization requirements input by the user, wherein the cultural and creative product customization requirements include aesthetic preferences, cultural background and description of specific occasion requirements; S120, performing image semantic feature extraction on the product prototype image to obtain a fine-grained semantic coding feature map of the product prototype image; S130, performing standardized expression and semantic coding on the cultural and creative product customization requirements to obtain a semantic coding feature vector of the cultural and creative product customization requirements; S140, based on the semantic coding feature vector of the cultural and creative product customization requirements, performing fine-grained semantic interactive optimization modulation on the fine-grained semantic coding feature map of the product prototype image to obtain a customization requirement modulated product image feature inversion feature map; S150, performing customization effect generation based on the customization requirement modulated product image feature inversion feature map to obtain a customized effect image of the cultural and creative product.

[0055] In step S110, the product prototype image is extracted from the product prototype library, and the customization requirements of the cultural and creative products input by the user are obtained at the same time. The customization requirements of the cultural and creative products include aesthetic preferences, cultural background and descriptions of needs for specific occasions. Among them, the user's aesthetic preferences, cultural background and needs for specific occasions are key factors that constitute the personalized characteristics of cultural and creative products. For example, the user may expect that specific cultural elements can be incorporated into the product design, or that it conforms to a certain specific aesthetic style, or that it is a special style designed for a specific occasion (such as a festival, celebration, etc.). By obtaining the customization requirements input by the user, it is helpful to accurately understand the user's expectations, thereby generating customized cultural and creative products that better meet the user's personalized needs. In addition, as the basis for creativity and design, the product prototype library provides a series of pre-designed basic templates or examples. By selecting the prototype closest to the customer's needs from the library as the design starting point, it can respond to the user's customization needs more quickly, thereby saving time and resources in creating a design from scratch.

[0056] Specifically, the system first maintains a library of product prototypes containing a variety of basic design elements and styles. These prototypes can be completed design templates, basic shapes, or color schemes, aiming to provide users with a diverse base of choices. When users access the customization service, they can browse this library and select the product that most closely resembles their imagined design as a starting point. This step not only helps users quickly identify a design direction that meets their general conception, but also provides a good starting point for subsequent, more refined adjustments.

[0057] Next, after users have selected a satisfactory product prototype, the system will guide them to a detailed customization requirement input interface. Users are encouraged to describe in detail their aesthetic preferences, cultural background, and the needs of any specific occasion. For example, if they are preparing a gift to celebrate a festival or anniversary, they may particularly emphasize that they hope the finished product can reflect the relevant theme; or if they want a product that can represent their personal style and taste, then specific requirements regarding color matching, pattern selection, and even material texture are particularly important. By providing as much detailed information as possible, we can better understand user expectations and make targeted design adjustments accordingly.

[0058] In order to ensure that the information collection process is both efficient and comprehensive, the system usually uses a structured approach to guide users to fill in relevant information. For example, a series of preset questions are set up to allow users to choose appropriate options based on their own circumstances, while leaving open text boxes for additional explanations. This can not only ensure that the collected data has a certain degree of standardization, which is convenient for subsequent processing and analysis, but also fully respect the unique ideas of each customer. In addition, considering the possible differences in expression ability among different users, the system may also introduce natural language processing technology to assist in parsing unstructured text content, further improving the accuracy of understanding user intentions. In short, by effectively utilizing the product prototype library and accurately capturing user customization needs, it is possible to quickly respond to market changes while maintaining creative diversity and meet the growing demand for personalized consumption.

[0059] In step S120, image semantic feature extraction is performed on the product prototype image to obtain a fine-grained semantic coding feature map of the product prototype image. Specifically, in an embodiment of the present application, image semantic feature extraction is performed on the product prototype image to obtain a fine-grained semantic coding feature map of the product prototype image, including: inputting the product prototype image into a product prototype image feature extractor based on the ViT model to obtain the fine-grained semantic coding feature map of the product prototype image. Accordingly, considering that the design of cultural and creative products involves not only tiny visual elements, but also the overall design style and product structure. Therefore, in order to improve the ability to understand the semantics of product prototype images, the present application adopts the ViT model to construct a product prototype image feature extractor, and performs fine-grained image semantic analysis on the product prototype image to obtain a fine-grained semantic coding feature map of the product prototype image. Those skilled in the art should be aware that, compared with traditional convolutional neural networks, the ViT model has significant advantages in processing long-distance dependencies and global contextual information. It divides the product prototype image into multiple small patches, linearly embeds and positionally encodes each small patch, and then uses a self-attention mechanism to capture the semantic dependencies between image patches. This allows it to effectively capture the global information and detailed features in the product prototype image, extract key design element features and deep semantic information in the product prototype image, and thus provide a solid feature foundation for subsequent customized feature modulation.

[0060] Specifically, the Vision Transformer (ViT) is a model based on the Transformer architecture that directly inputs images as a series of patches into the Transformer for processing. The introduction of the ViT model marks the beginning of the successful application of the Transformer architecture in the field of natural language processing to computer vision tasks, and has achieved performance that even surpasses traditional convolutional neural networks (CNNs) on multiple benchmarks. The following is a detailed description of the process of inputting a product prototype image into a product prototype image feature extractor based on the ViT model:

[0061] In the ViT model, the product prototype image must first be preprocessed. Assuming there is an RGB color image of size 224x224 pixels as input, in order to convert it into a form suitable for Transformer processing, the image is usually divided into small blocks of fixed size or patches. For example, 16x16 pixels can be selected as the size of a patch, so that the entire image is divided into 14x14=196 such small blocks. Each patch is then flattened into a vector. If the original image is in RGB format, each patch can be flattened into a one-dimensional vector of length 16*16*3=768.

[0062] Next, all of these flattened patch vectors are concatenated to form a sequence, which serves as the input to the Transformer model. Specifically, a special classification token [CLS] is added to the front of the sequence, similar to the [CLS] tag in NLP tasks, summarizing information about the entire image. Furthermore, to provide the model with positional information, a learnable positional encoding is added to each patch vector and [CLS] token. This positional encoding is trained to enable the model to understand the relative positional relationships between different patches, which is crucial for processing spatially structured data such as images.

[0063] After completing the above preparations, a tensor with a shape of (197,768) is obtained, where 197 represents the total number of patches including the [CLS] token, and 768 is the dimension of each patch vector. Next, this tensor will be fed into a multi-layer Transformer encoder. Each Transformer encoder layer consists of two sublayers: the first sublayer is a multi-head self-attention mechanism used to capture long-distance dependencies between different patches; the second sublayer is a feedforward network used to further process and enhance feature expression capabilities. Residual connections and layer normalization are applied between these two sublayers and after the entire encoder layer to ensure stable gradient propagation and accelerate the convergence process.

[0064] A typical ViT model may contain multiple such Transformer encoder layers stacked together, and the specific number of layers depends on the needs of the actual application scenario. After all encoder layers are processed, the tensor of shape (196,768) is reshaped into a form that is closer to the spatial structure of the original image. For example, the original image is 224x224 pixels and is divided into 14x14 patches. Therefore, each 768-dimensional feature vector can be regarded as a 14x14x768 three-dimensional tensor, which is equivalent to rearranging the features of each patch into a two-dimensional grid to restore some of the spatial structure information of the original image.

[0065] In summary, the ViT model effectively achieves the goal of extracting rich semantic information from image data by decomposing the image into a series of ordered patches and leveraging the powerful Transformer architecture to handle the complex relationships between these patches.

[0066] In step S130, the cultural and creative product customization requirements are standardized and semantically encoded to obtain a semantic encoding feature vector of the cultural and creative product customization requirements. Figure 2 This is a flow chart of the method for intelligent service of cultural and creative product customization according to the embodiment of the present application for standardizing the expression and semantic encoding of the cultural and creative product customization requirements to obtain the semantic encoding feature vector of the cultural and creative product customization requirements. Specifically, in the embodiment of the present application, if Figure 2 As shown, the customization requirements of cultural and creative products are standardizedly expressed and semantically encoded to obtain a semantically encoded feature vector of the customization requirements of cultural and creative products, including: S210, inputting the customization requirements of cultural and creative products into a customization requirement standard expression module based on a large language model to obtain a standardized expression of the customization requirements of cultural and creative products; S220, semantically encoding the standardized expression of the customization requirements of cultural and creative products to obtain a semantically encoded feature vector of the customization requirements of cultural and creative products.

[0067] Specifically, in step 210, the cultural and creative product customization requirements are input into a customization requirement specification expression module based on a large language model to obtain a cultural and creative product customization requirement specification expression. In particular, it is considered that the cultural and creative product customization requirements input by the user may be unstructured, vague, or use colloquial or personalized language descriptions. Therefore, in order to reduce ambiguity and ensure the accurate expression of customization requirements, the present application further uses a customization requirement specification expression module based on a large language model to standardize the cultural and creative product customization requirements, and convert them into a standard, clear, easy-to-understand and easy-to-execute language form. In a specific example of the present application, the large language model is a GPT-3 (Generative Pre-trained Transformer 3) model, which has powerful language understanding and generation capabilities through large-scale pre-training, can fully understand the semantic meaning of the cultural and creative product customization requirements, and convert them into clear and standardized customization requirement descriptions, generate cultural and creative product customization requirement specification expressions, and ensure the accuracy of the subsequent cultural and creative product customization process.

[0068] Specifically, in step S220, the standardized expression of cultural and creative product customization requirements is semantically encoded to obtain a semantically encoded feature vector of the cultural and creative product customization requirements. Accordingly, in order to convert the textual description of product customization requirements into a form understandable by the machine learning model, this application further uses natural language processing technology to semantically encode the standardized expression of cultural and creative product customization requirements, mapping it to a high-dimensional semantic feature space, and deeply understanding its contextual meaning, thereby generating a semantically encoded feature vector of cultural and creative product customization requirements. This enables the standardized expression of cultural and creative product customization requirements to effectively interact with product prototype images across modal semantics. In a specific example of this application, the BERT model is used to perform the above-mentioned semantic encoding process.

[0069] Among them, in the process of customizing cultural and creative products, using the BERT model to semantically encode the cultural and creative product customization requirements input by users is one of the key steps to achieve efficient and accurate understanding of user intentions. The BERT model is known for its powerful two-way context understanding capabilities and has a wide range of applications in the field of natural language processing, including text classification, sentiment analysis, question-answering systems, etc. When applied to the customization of cultural and creative products, BERT can help the system deeply analyze the user's text description and convert it into a form that is easy for computers to process and rich in semantic information - that is, the semantic encoding feature vector of the cultural and creative product customization requirements. This process not only improves the efficiency and quality of the subsequent design stages, but also ensures that the final product is closer to the real needs of users.

[0070] Specifically, the process of using the BERT model to semantically encode the standard expression of the cultural and creative product customization requirements is as follows:

[0071] First, users submit customization requirements for cultural and creative products through a specific interface. These requirements usually include aesthetic preferences, cultural background, and descriptions of needs for specific occasions. In order for the BERT model to work effectively, the original text data needs to be preprocessed. This step mainly includes word segmentation (Tokenization), adding special tags (such as [CLS] and [SEP]), and converting the text into a format acceptable to the model. Among them, word segmentation refers to splitting a sentence into individual words or subword units, and adding special tags is to let the model know the start and end positions of a sequence. In addition, since BERT has a fixed input length limit, if the text is too long, it may need to be truncated or split.

[0072] After the aforementioned preprocessing, the text data is fed into the pre-trained BERT model. Here, BERT utilizes its multi-layer Transformer encoder structure to capture the contextual meaning of each word within the entire sentence. Specifically, each encoder layer consists of two main components: a self-attention mechanism and a feedforward neural network. The self-attention mechanism allows the model to simultaneously focus on information at different positions within the sentence and dynamically adjust weights based on the relationships between them; the feedforward network further enhances the feature representation capabilities. The entire encoding process is bidirectional, meaning that the model can learn the meaning of the text from both left-to-right and right-to-left directions, thereby gaining a more comprehensive understanding.

[0073] As data flows through each encoder layer, each word gradually forms a high-dimensional vector representation that comprehensively reflects all relevant information about the word and its context. For the entire sentence, the vector corresponding to the [CLS] token at the beginning of the sequence is particularly important because it embodies the main meaning of the entire sentence and is often used as input for downstream tasks. However, in the context of customized cultural and creative products, the greater concern is how to translate the user's specific needs into a series of actionable design guidelines. Therefore, in addition to [CLS], the semantic encoding results of other key words or phrases need to be considered.

[0074] To achieve this, a common practice is to extract the hidden states of all words from the last layer of BERT output as preliminary semantic encoding features. Then, based on the actual application requirements, different strategies can be used to further refine these features. For example, the vectors of related words can be aggregated by weighted averaging, maximum pooling, etc. to generate a single feature vector representing the entire customization requirement. Another method is to directly select those key words related to the design (such as color, pattern style, etc.) and use their respective vectors to combine into the final feature representation. In this way, the system can quickly and accurately understand the user's needs and provide them with a more personalized service experience.

[0075] In step S140, based on the semantic encoding feature vector of the cultural and creative product customization requirements, the fine-grained semantic encoding feature map of the product prototype image is subjected to fine-grained semantic interaction optimization modulation to obtain a customization requirement modulated product image feature inversion feature map. Furthermore, in order to effectively utilize the semantic information of the cultural and creative product customization requirements to guide the customized feature modulation of the product prototype image, the present application proposes a cross-modal semantic interaction method based on prompt information. By using the semantic encoding feature vector of the cultural and creative product customization requirements as prompt information, the fine-grained semantic interaction optimization modulation of the fine-grained semantic encoding feature map of the product prototype image is performed to achieve deep cross-modal semantic fusion, so that the customized modulated product image features can fully reflect the user's personalized needs.

[0076] Figure 3 This is a flow chart of performing fine-grained semantic interaction optimization modulation on the fine-grained semantic coding feature map of the product prototype image based on the semantic coding feature vector of the cultural and creative product customization demand in the intelligent service method for cultural and creative product customization according to the embodiment of the present application to obtain a customized demand modulation product image feature inversion feature map. Specifically, in the embodiment of the present application, Figure 3As shown, based on the semantic coding feature vector of the cultural and creative product customization demand, the fine-grained semantic coding feature map of the product prototype image is subjected to fine-grained semantic interactive optimization modulation to obtain a customization demand modulated product image feature inversion feature map, including: S310, autocorrelation coding is performed on the semantic coding feature vector of the cultural and creative product customization demand to obtain a cultural and creative product customization demand autocorrelation coding matrix; S320, local fine-grained semantic interactive coding is performed on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices; S330, based on the set of customization demand-product prototype image cross-modal prompt information semantic coding matrices, the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map are guided to perform cross-modal interactive optimization to obtain the customization demand modulated product image feature inversion feature map.

[0077] Specifically, in step S310, the semantically encoded feature vector of the cultural and creative product customization requirements is subjected to autocorrelation encoding to obtain an autocorrelation encoding matrix of the cultural and creative product customization requirements. This includes calculating the product between the semantically encoded feature vector of the cultural and creative product customization requirements and the transposed vector of the semantically encoded feature vector of the cultural and creative product customization requirements to obtain the autocorrelation encoding matrix of the cultural and creative product customization requirements. In other words, by calculating the product between the semantically encoded feature vector of the cultural and creative product customization requirements and its transposed vector, the concept of outer product in matrix operations is utilized to reveal the autocorrelation of the cultural and creative product customization requirements, thereby obtaining the autocorrelation encoding matrix of the cultural and creative product customization requirements.

[0078] The processing of step S310 can be expressed by the following formula:

[0079]

[0080] Among them, V1 represents the semantic encoding feature vector of the cultural and creative product customization demand, represents the matrix multiplication operation, (·) T Represents the transpose of a vector, M z Represents the autocorrelation coding matrix of cultural and creative product customization demand.

[0081] More specifically, in step S320, local fine-grained semantic interaction encoding is performed on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices, including: feature decoupling of the product prototype image fine-grained semantic coding feature map to obtain a set of product prototype image local semantic feature matrices; linear transformation is performed on the cultural and creative product customization demand autocorrelation coding matrix to obtain a cultural and creative product customization demand query coding matrix and a cultural and creative product customization demand value coding matrix; each product prototype image local semantic feature matrix in the set of the cultural and creative product customization demand query coding matrix, the cultural and creative product customization demand value coding matrix and the product prototype image local semantic feature matrix is respectively input into a cross-modal prompt information encoder based on a converter structure to obtain the set of customization demand-product prototype image cross-modal prompt information semantic coding matrices.

[0082] That is, by performing feature decoupling on the fine-grained semantic coding feature map of the product prototype image, it is decomposed into multiple local semantic feature matrices of the product prototype image to improve the model's attention to the local details of the product prototype image and enhance the distinguishability of the features, thereby facilitating the execution of a more fine-grained cross-modal interaction analysis. Subsequently, the cultural and creative product customization demand autocorrelation coding matrix is converted into a cultural and creative product customization demand query coding matrix and a cultural and creative product customization demand value coding matrix through linear mapping, so that they meet the input form requirements of the attention mechanism. Among them, the cultural and creative product customization demand query coding matrix is used to query the correlation features with the product prototype image, while the cultural and creative product customization demand value coding matrix stores the original cultural and creative product customization demand information to be fused. Furthermore, based on the converter structure, the cultural and creative product customization demand query encoding matrix, the cultural and creative product customization demand value encoding matrix and the local semantic feature matrix of each product prototype image are cross-modally encoded. Through the self-attention mechanism of the converter structure, the feature fusion and interactive analysis of cross-modal data are realized, and the deep semantic connection between the cultural and creative product customization demand and the product prototype image features is effectively captured. A set of customization demand-product prototype image cross-modal prompt information semantic encoding matrices is generated, and on this basis, the subsequent feature interaction optimization is further guided.

[0083] The processing of step S320 can be expressed by the following formula:

[0084] Decouple(X)={M1,M2,...,M i ,...M n}

[0085]

[0086] Where X represents the fine-grained semantic encoding feature map of the product prototype image, Decouple(·) represents feature decoupling, M1, M2, M i and M n They represent the first, second, i-th and n-th local semantic feature matrices of product prototype images in the set of local semantic feature matrices of product prototype images, respectively. n is the number of local semantic feature matrices of product prototype images. W q and W v Denote the query embedding matrix and the value embedding matrix respectively, M q and M v They represent the cultural and creative product customization demand query coding matrix and the cultural and creative product customization demand value coding matrix respectively, M i T It's M i The transposed matrix of , S is the characteristic scale value of the local semantic feature matrix of the product prototype image, softmax(·) represents the normalized exponential function, I i Represents the semantic encoding matrix of the cross-modal prompt information of the i-th customization requirement-product prototype image.

[0087] More specifically, in step S330, based on the set of semantic coding matrices of cross-modal prompt information of customized demand-product prototype image, the semantic coding feature vector of customized demand of cultural and creative products and the fine-grained semantic coding feature map of product prototype image are guided to perform cross-modal interaction optimization to obtain the feature inversion feature map of customized demand modulated product image, including: inputting the set of semantic coding matrices of cross-modal prompt information of customized demand-product prototype image into the information gating unit based on the decoder to obtain the set of cross-modal semantic interaction attention weights of customized demand-product prototype image; inputting the set of autocorrelation coding matrix of cultural and creative product customized demand and the set of local semantic feature matrix of product prototype image into the cross-modal interaction unit to obtain the set of cross-modal interaction local feature matrix of customized demand-product prototype image; inputting the set of cross-modal semantic interaction attention weights of customized demand-product prototype image and the set of cross-modal interaction local feature matrix of customized demand-product prototype image into the cross-modal interaction optimization unit to obtain the feature inversion feature map of customized demand modulated product image.

[0088] Specifically, the set of semantic encoding matrices for cross-modal cue information between customization requirements and product prototype images is further input into a decoder-based information gating unit for information screening. The information gating unit uses the decoder to learn the feature importance of each semantic encoding matrix for cross-modal cue information between customization requirements and product prototype images, and accordingly assigns weights to generate a set of attention weights for cross-modal semantic interactions between customization requirements and product prototype images. This allows the important semantic connections between product prototype images and cultural and creative product customization requirements to be effectively highlighted during subsequent feature interaction, while suppressing interference from irrelevant information. Subsequently, a direct interaction is performed between the autocorrelation encoding matrix for cultural and creative product customization requirements and the set of local semantic feature matrices for product prototype images via a positional dot product method, thereby capturing the association information between the two at the same feature space location. The attention weights generated in this process are then used to weight and aggregate the interaction features between cultural and creative product customization requirements and product prototype images, thereby enhancing the expressive power of the features and restoring the original image structure. This generates a product image feature representation that incorporates user personalized requirement information, namely, the inversion feature map of the product image feature modulated by customization requirements.

[0089] The processing of step S330 can be expressed as follows:

[0090] a i =softmax{decoder(I i ,W / )}

[0091] f0=couple{a1·M1⊙M z ,a2·M2⊙M z ,...,a n ·M n ⊙M z}

[0092] Among them, I i represents the semantic encoding matrix of the cross-modal prompt information of the i-th customization requirement-product prototype image, decoder(·) represents the decoder, W / represents the decoding weight matrix, softmax(·) represents the normalized exponential function, a1, a2, a i and a n represents the attention weights of the cross-modal semantic interaction between the first, second, i-th and n-th customization requirement and product prototype image, ⊙ represents the positional dot product, couple{·} represents feature concatenation, and f0 represents the feature inversion map of the customization requirement modulated product image features.

[0093] In step S150, based on the customization requirement, the product image feature inversion feature map is modulated to perform customization effect generation to obtain a customized effect image of a cultural and creative product. Specifically, in an embodiment of the present application, based on the customization requirement, the product image feature inversion feature map is modulated to perform customization effect generation to obtain a customized effect image of a cultural and creative product, including: inputting the customization requirement modulated product image feature inversion feature map into a customization effect generator based on a diffusion model to obtain the customized effect image of the cultural and creative product. Among them, the diffusion model is a technology that has made significant progress in the field of image generation in recent years. Its working principle is similar to the diffusion process in the physical world: starting from an initial state, it gradually evolves to a more complex state until the final goal is reached. In the image generation task of the present application, the diffusion model receives the customization requirement modulated product image feature inversion feature map as input, and through a series of preset diffusion steps, gradually increases the details and structure of the image until a clear cultural and creative product image that highly matches the user's customization requirements is formed.

[0094] Specifically, the specific processing process of inputting the customized demand modulated product image feature inversion feature map into the customized effect generator based on the diffusion model is as follows:

[0095] First, the inverted feature map of the product image features modulated by the customization requirements described above is prepared as a conditional input. This means that during the generation process, the diffusion model not only considers how to construct beautiful and meaningful images, but also pays special attention to ensuring that these images are closely centered around the user's customization requirements. To this end, the diffusion model incorporates a series of carefully designed network layers, including but not limited to encoder-decoder architectures, residual connections, and attention mechanisms. These components work together to enable the model to incorporate new creative elements while retaining the essence of the original design.

[0096] During this process, the diffusion model uses a phased approach to image generation. Initially, the model may generate a rough outline or framework based on the provided feature map. This is primarily used to determine the overall layout and general style. As the number of iterations increases, the model continuously refines various aspects of the image, such as color matching, texture style, and pattern design, making it richer and more detailed. Specifically, at each stage, the model adjusts the generation strategy by referencing the inverted feature map of the product image features modulated by custom requirements, ensuring that each step moves in the direction of user expectations.

[0097] To improve generation quality, the diffusion model also employs advanced techniques such as adaptive sampling and multi-scale fusion. Adaptive sampling allows the model to dynamically adjust parameter settings in subsequent steps based on the quality of the currently generated image, thus avoiding falling into local optimal solutions. Multi-scale fusion, by simultaneously considering image information at different resolutions, ensures a balance between detailed representation and overall coordination in the final output.

[0098] Throughout the generation process, the system can also periodically display intermediate results to the user, allowing for timely feedback and adjustments. This interactive design approach not only improves user experience satisfaction but also promotes continuous optimization and refinement of creative ideas. Ultimately, when the model deems it has reached a satisfactory level, it stops iteration and outputs the finished product: a rendering of a cultural and creative product tailored entirely to the user's needs. This approach not only greatly enriches the expression of cultural and creative products but also provides consumers with unprecedented personalized choices.

[0099] In particular, here, when the product prototype image fine-grained semantic coding feature map and the cultural and creative product customization demand semantic coding feature vector respectively represent the image fine-grained semantic coding features of the product prototype image and the textual semantic coding features expressed by the cultural and creative product customization demand specification, considering the auxiliary differences in their cross-modal prompts, it is expected to improve the semantically consistent aggregation expression effect of the customization demand modulated product image feature inversion feature map obtained through cross-modal interaction optimization.

[0100] Preferably, inputting the customized demand modulated product image feature inversion feature map into a customized effect generator based on a diffusion model to obtain the customized effect image of the cultural and creative product includes:

[0101] Perform feature clustering on the feature set of the customized demand modulated product image feature inversion feature map to obtain a customized demand modulated product image feature inversion class feature set and a customized demand modulated product image feature inversion class feature set, namely:

[0102]

[0103] in, is the feature set within the product image feature inversion class modulated by the customization requirement, f 1i is the feature value of each position in the feature set of the custom demand modulation product image feature inversion class, f 24 is the feature value of each position in the out-of-class feature set of the customized demand modulation product image feature inversion;

[0104] Calculate the ratio of the number of eigenvalues in the feature set of the custom demand modulated product image feature inversion class to the number of eigenvalues in the feature set of the custom demand modulated product image feature inversion feature map, that is:

[0105]

[0106] Wherein, k is the number of eigenvalues in the feature set of the custom demand modulated product image feature inversion class, m is the number of eigenvalues in the feature set of the custom demand modulated product image feature inversion feature map, and λ is the ratio of k to m;

[0107] The ratio of the λth power of the sum of the absolute values of all eigenvalues in the feature set of the customized demand modulated product image feature inversion class to the λth power of the sum of the absolute values of all eigenvalues in the feature set of the customized demand modulated product image feature inversion feature map is calculated to obtain the customized demand modulated product image feature inversion modulation weight, that is:

[0108]

[0109] Wherein, w1 is the inversion modulation weight of the image feature of the customized demand modulation product;

[0110] Calculate the ratio of the λ / 2 power of the sum of the squares of all eigenvalues in the feature set of the customized demand modulated product image feature inversion class to the λ / 2 power of the sum of the squares of all eigenvalues in the feature set of the customized demand modulated product image feature inversion feature map to obtain the customized demand modulated product image feature inversion harmonic weight, that is:

[0111]

[0112] Wherein, w2 is the inversion and harmonic weight of the image feature of the customized demand modulation product;

[0113] For each eigenvalue in the feature set within the customized demand modulated product image feature inversion class, calculate the product of the eigenvalue and the customized demand modulated product image feature inversion harmonic weight, and then add the customized demand modulated product image feature inversion modulation weight to obtain the optimized eigenvalue, that is:

[0114] f′ 1i =w2×f 1i +w1

[0115] Among them, f′ 1i is the optimized feature value of each position of the feature set within the image feature inversion class of the customized demand modulation product;

[0116] For each eigenvalue in the out-of-class feature set of the customized demand modulated product image feature inversion, the product of the eigenvalue and the weight of the customized demand modulated product image feature inversion modulation value is calculated to obtain the optimized eigenvalue, that is:

[0117] f 24 ′=w1×f 24

[0118] Among them, f 24 ' is the optimized feature value of each position of the out-of-class feature set of the customized demand modulation product image feature inversion;

[0119] The optimized customized demand modulated product image feature inversion feature map composed of the optimized feature values of the feature set within the customized demand modulated product image feature inversion class and the feature set outside the customized demand modulated product image feature inversion class is input into the customized effect generator based on the diffusion model to obtain the customized effect image of the cultural and creative product.

[0120] Therefore, while clustering the feature inversion feature map of the customized demand modulated product image, the interactive description of the key feature information of the customized demand modulated product image inversion feature map in the clustering process is carried out, and the geometric equivariant topology of the feature is constructed by low-rank harmonic modulation based on the equivariance of the clustering features and the feature as a whole of the customized demand modulated product image feature inversion feature map, so as to obtain the graphical distribution translation and rotation symmetry of the clustering features of the customized demand modulated product image feature inversion feature map relative to the feature as a whole, thereby introducing geometric message passing into the feature expression of the customized demand modulated product image feature inversion feature map and realizing the cluster mapping symmetry of the customized demand modulated product image feature inversion feature map through irreducible low-rank order coefficient manipulation, thereby improving the clustering-based feature representation consistency of the customized demand modulated product image feature inversion feature map, thereby improving the image quality of the customized effect image of the cultural and creative product obtained by inputting the customized effect generator based on the diffusion model into the customized demand modulated product image feature inversion feature map.

[0121] In summary, the intelligent service method for customized cultural and creative products based on the embodiment of the present application is explained. It extracts product prototype images from the product prototype library as initial materials, and uses deep learning-based image processing technology to perform semantic analysis on the prototype images of cultural and creative products, thereby capturing the key design elements in the images. At the same time, combined with the cultural and creative product customization requirements input by the user, through fine-grained cross-modal semantic interaction between the product customization requirements and the product prototype images, customized feature modulation of the product prototype images is achieved according to the product customization requirements, and customized effect images of cultural and creative products are intelligently generated based on this. In this way, it is possible to quickly respond to users' personalized customization needs, shorten the design cycle, improve customization efficiency, reduce customization costs, and effectively enhance the user's customization experience.

[0122] Figure 4 This is a block diagram of the intelligent service system for customizing cultural and creative products according to an embodiment of the present application. Figure 4 As shown, the intelligent service system 100 for customization of cultural and creative products according to the embodiment of the present application includes: a product image customization demand acquisition module 110, which is used to extract product prototype images from a product prototype library and simultaneously obtain cultural and creative product customization demands input by a user, wherein the cultural and creative product customization demands include aesthetic preferences, cultural background and descriptions of demands for specific occasions; a product prototype image semantic feature extraction module 120, which is used to perform image semantic feature extraction on the product prototype image to obtain a fine-grained semantic coding feature map of the product prototype image; a customization demand semantic coding module 130, which is used to perform standardized expression and semantic coding on the cultural and creative product customization demands to obtain a semantic coding feature vector of the cultural and creative product customization demands; a customization demand modulation module 140, which is used to perform fine-grained semantic interactive optimization modulation on the fine-grained semantic coding feature map of the product prototype image based on the semantic coding feature vector of the cultural and creative product customization demands to obtain a customization demand modulated product image feature inversion feature map; and a customization effect generation module 150, which is used to generate customization effects based on the customization demand modulated product image feature inversion feature map to obtain a customized effect image of the cultural and creative product.

[0123] The specific operations of each step in the above-mentioned cultural and creative product customization intelligent service system have been referenced above. Figures 1 to 3 The intelligent service method for customization of cultural and creative products has been introduced in detail, and therefore, its repeated description will be omitted.

[0124] As described above, the intelligent service system 100 for customizing cultural and creative products according to the embodiments of the present disclosure can be implemented in various wireless terminals, such as a server equipped with an intelligent service algorithm for customizing cultural and creative products. In one possible implementation, the intelligent service system 100 for customizing cultural and creative products according to the embodiments of the present disclosure can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the intelligent service system 100 for customizing cultural and creative products can be a software module in the operating system of the wireless terminal, or an application developed for the wireless terminal; of course, the intelligent service system 100 for customizing cultural and creative products can also be one of the many hardware modules of the wireless terminal.

[0125] Alternatively, in another example, the cultural and creative product customization intelligent service system 100 and the wireless terminal may also be separate devices, and the cultural and creative product customization intelligent service system 100 may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.

[0126] The foregoing is merely an example of the principles of the present disclosure, and various modifications may be made by those skilled in the art without departing from the scope of the present disclosure. The above embodiments are presented for the purpose of illustration and not limitation. The present disclosure may also take many forms other than those explicitly described herein.

Claims

1. An intelligent service method for customizing cultural and creative products, characterized in that: include: Extracting product prototype images from a product prototype library and simultaneously obtaining cultural and creative product customization requirements input by users, wherein the cultural and creative product customization requirements include aesthetic preferences, cultural background, and occasion requirements description; Performing image semantic feature extraction on the product prototype image to obtain a fine-grained semantic encoding feature map of the product prototype image; Standardizing and semantically encoding the cultural and creative product customization requirements to obtain a semantic encoding feature vector of the cultural and creative product customization requirements; Based on the cultural and creative product customization demand semantic coding feature vector, fine-grained semantic interaction optimization modulation is performed on the product prototype image fine-grained semantic coding feature map to obtain a customization demand modulation product image feature inversion feature map, which includes: Performing autocorrelation coding on the semantic coding feature vector of the cultural and creative product customization demand to obtain an autocorrelation coding matrix of the cultural and creative product customization demand; Performing local fine-grained semantic interaction coding on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices; Based on the set of the customization demand-product prototype image cross-modal prompt information semantic coding matrices, guiding the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to perform cross-modal interactive optimization to obtain the customization demand modulated product image feature inversion feature map; Based on the customization requirements, the product image feature inversion feature map is modulated to generate a customization effect to obtain a customized effect image of the cultural and creative product.

2. The intelligent service method for customizing cultural and creative products according to claim 1, characterized in that: Performing image semantic feature extraction on the product prototype image to obtain a fine-grained semantic encoding feature map of the product prototype image includes: The product prototype image is input into a product prototype image feature extractor based on the ViT model to obtain a fine-grained semantic encoding feature map of the product prototype image.

3. The intelligent service method for customizing cultural and creative products according to claim 2, characterized in that: The cultural and creative product customization requirements are standardized and semantically encoded to obtain a semantic encoding feature vector of the cultural and creative product customization requirements, including: Inputting the cultural and creative product customization requirements into a customization requirement specification expression module based on a large language model to obtain a cultural and creative product customization requirement specification expression; The cultural and creative product customization requirement specification expression is semantically encoded to obtain the cultural and creative product customization requirement semantic encoding feature vector.

4. The intelligent service method for customizing cultural and creative products according to claim 3, characterized in that: Performing autocorrelation coding on the semantic coding feature vector of the cultural and creative product customization demand to obtain an autocorrelation coding matrix of the cultural and creative product customization demand, including: The product between the semantic coding feature vector of the cultural and creative product customization demand and the transposed vector of the semantic coding feature vector of the cultural and creative product customization demand is calculated to obtain the autocorrelation coding matrix of the cultural and creative product customization demand.

5. The intelligent service method for customizing cultural and creative products according to claim 4, characterized in that: Performing local fine-grained semantic interaction coding on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices, including: Performing feature decoupling on the fine-grained semantic encoding feature map of the product prototype image to obtain a set of local semantic feature matrices of the product prototype image; Performing a linear transformation on the cultural and creative product customization demand autocorrelation coding matrix to obtain a cultural and creative product customization demand query coding matrix and a cultural and creative product customization demand value coding matrix; The local semantic feature matrices of each product prototype image in the set of the cultural and creative product customization demand query coding matrix, the cultural and creative product customization demand value coding matrix and the product prototype image local semantic feature matrix are respectively input into the cross-modal prompt information encoder based on the converter structure to obtain the set of the customization demand-product prototype image cross-modal prompt information semantic coding matrices.

6. The intelligent service method for customizing cultural and creative products according to claim 5, characterized in that: Based on the set of the semantic coding matrix of the cross-modal prompt information of the customization demand and the product prototype image, guiding the cross-modal interactive optimization of the semantic coding feature vector of the customization demand of the cultural and creative product and the fine-grained semantic coding feature map of the product prototype image to obtain the feature inversion feature map of the customization demand modulated product image, including: Inputting the set of the customization demand-product prototype image cross-modal prompt information semantic encoding matrices into the decoder-based information gating unit to obtain a set of customization demand-product prototype image cross-modal semantic interaction attention weights; Inputting the set of the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image local semantic feature matrix into a cross-modal interaction unit to obtain a set of customization demand-product prototype image cross-modal interaction local feature matrices; The set of attention weights of the cross-modal semantic interaction between the customization requirement and the product prototype image and the set of local feature matrices of the cross-modal interaction between the customization requirement and the product prototype image are input into the cross-modal interaction optimization unit to obtain the feature inversion feature map of the customization requirement modulated product image.

7. The intelligent service method for customizing cultural and creative products according to claim 6, characterized in that: Modulating the product image feature inversion feature map based on the customization requirement to generate a customization effect to obtain a customized effect image of the cultural and creative product, including: The customized demand modulated product image feature inversion feature map is input into a customized effect generator based on a diffusion model to obtain the customized effect image of the cultural and creative product.

8. An intelligent service system for customizing cultural and creative products, the system being configured to execute the method according to claim 7, characterized in that: include: A product image customization demand acquisition module is used to extract product prototype images from the product prototype library and simultaneously acquire cultural and creative product customization demands input by users, wherein the cultural and creative product customization demands include aesthetic preferences, cultural background, and occasion demand descriptions; A product prototype image semantic feature extraction module, configured to extract image semantic features from the product prototype image to obtain a fine-grained semantic encoding feature map of the product prototype image; A customization requirement semantic coding module, configured to perform standardized expression and semantic coding on the customization requirement of the cultural and creative product to obtain a semantic coding feature vector of the customization requirement of the cultural and creative product; A customization demand modulation module is used to perform fine-grained semantic interaction optimization modulation on the fine-grained semantic coding feature map of the product prototype image based on the semantic coding feature vector of the customization demand of the cultural and creative product to obtain a customization demand modulation product image feature inversion feature map, which includes: Performing autocorrelation coding on the semantic coding feature vector of the cultural and creative product customization demand to obtain an autocorrelation coding matrix of the cultural and creative product customization demand; Performing local fine-grained semantic interaction coding on the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to obtain a set of customization demand-product prototype image cross-modal prompt information semantic coding matrices; Based on the set of the customization demand-product prototype image cross-modal prompt information semantic coding matrices, guiding the cultural and creative product customization demand autocorrelation coding matrix and the product prototype image fine-grained semantic coding feature map to perform cross-modal interactive optimization to obtain the customization demand modulated product image feature inversion feature map; The customized effect generation module is used to modulate the product image feature inversion feature map based on the customized requirements to generate customized effects to obtain customized effect images of cultural and creative products.

9. The intelligent service system for customizing cultural and creative products according to claim 8, characterized in that: The customized effect generation module is used to: The customized demand modulated product image feature inversion feature map is input into a customized effect generator based on a diffusion model to obtain the customized effect image of the cultural and creative product.

Citation Information

Patent Citations

  • Commodity personalized customization method and platform based on artificial intelligence, and business model

    CN116883115A

  • Music generation method and training method and device of music generation model

    CN117496963A

Cited By

  • Method and system for evaluating customized design of cultural and creative products based on three-dimensional model of ancient building

    CN121303971A