Efficient image analysis and question-answering system for intelligent customer service
By optimizing visual information processing through multi-block word fusion and spatial word fusion technologies, the problems of slow response speed and high cost in intelligent customer service are solved, instant and high-precision question-and-answer capabilities are achieved, and the user experience and operational efficiency of the e-commerce platform are improved.
Patent Information
- Application Number
- CN202510821510.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-26
AI Technical Summary
Existing visual language models have slow response speeds and high operating costs in intelligent customer service scenarios, making it difficult to improve efficiency while ensuring answer quality.
Multi-block word fusion and spatial word fusion technologies are introduced to optimize the visual information processing process, reduce redundant data, shorten the amount of data processed by the language model, and combine two-stage training to optimize parameters.
It significantly improves response speed, reduces computing resource requirements, ensures instant and high-precision question-answering capabilities, and improves user experience and the operational efficiency of e-commerce platforms.
Smart Images

Figure CN120708027A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence applications, specifically to an efficient image analysis and question-answering system for intelligent customer service. This system aims to address the slow response speed and high operating costs of existing visual language models in practical applications. By optimizing the visual information processing process, it provides a fast, accurate, and cost-effective solution for intelligent customer service scenarios. Background Art
[0002] Advanced AI models capable of understanding images and engaging in conversations currently demonstrate enormous commercial potential for intelligent customer service, particularly in e-commerce. However, their scalable deployment faces significant challenges. In an ideal shopping scenario, a customer uploads a product image and immediately asks questions such as, "Does this shirt zipper on the side or back?" or "What is the specific model of the item in the picture?" A successful intelligent customer service representative must be able to answer these questions instantly and accurately to facilitate transactions. However, existing technologies face significant challenges in achieving this: image analysis requires extensive and complex computations, which not only consumes significant amounts of high-end computing resources, resulting in high costs per query, but also, more critically, unacceptable response delays. In the fast-paced online shopping world, any delay can cause customers to lose patience and abandon their purchases. While the industry has attempted to increase speed by simplifying models, this often sacrifices analysis accuracy and detail recognition capabilities, resulting in irrelevant answers or inability to clearly discern details, severely impacting user experience and brand reputation. Therefore, ensuring response quality while overcoming efficiency and cost bottlenecks has become a key obstacle in transforming advanced visual AI into a truly efficient and reliable intelligent customer service tool. Summary of the Invention
[0003] The present invention provides an efficient and low-cost image analysis and question-answering system for intelligent customer service, aiming to overcome the efficiency bottleneck of existing technologies in intelligent customer service. This system introduces an innovative visual information compression engine, which can efficiently refine and condense image information before passing it to the core language agent, remove redundant data, and retain the most critical visual features. This processing method significantly reduces the amount of data that needs to be processed by the subsequent language model, thereby increasing the response speed of the entire question-answering process by orders of magnitude, ensuring that customers can get almost instantaneous answers after asking questions. The final result is a practical system that can empower intelligent customer service. It can quickly respond to any questions from customers about product images, maintain a high degree of analysis accuracy, and accurately identify product details. This not only greatly improves the user shopping experience, but also provides strong technical support for e-commerce platforms to improve customer satisfaction and sales conversion rates.
[0004] To achieve the above object, the technical solution of the present invention is as follows:
[0005] An efficient image analysis and question-answering system for intelligent customer service, including the following steps:
[0006] Step 1: Perform Multi-Block Token Fusion (MBTF) to integrate multi-granular features from different levels of the visual encoder.
[0007] Step 2: Perform spatial token fusion (STF) on the fused word sequence or the original visual word sequence in step 1 to reduce the spatial redundancy of visual words.
[0008] Step 3: Align the compact visual tokens generated in step 2 with the text embedding space of the large-scale language model.
[0009] Step 4: Use the aligned compact visual tokens for model inference and optimize the learnable parameters involved in the entire method through a specific two-stage training process.
[0010] Furthermore, in step 1, the multi-block word fusion MBTF is designed to enable the model to adapt to the needs of a wide range of downstream visual-language tasks for different granularity features and to supplement the detailed information that may be lost due to word reduction. This step extracts visual features from multiple different levels of the visual encoder. Specifically, from M pre-selected levels of the visual encoder (For example, for a ViT-L / 14 encoder with 24 blocks, 8 levels can be uniformly sampled) Extract the respective visual word units Then, the extracted visual words Reshape into a feature map with predetermined spatial and channel dimensions. Next, the reshaped feature maps are concatenated along their channel dimensions. Finally, the concatenated features are processed using at least one convolution operation with a kernel size of 1×1 followed by an activation function (e.g., GeLU) to fuse multi-level features and gradually adjust the channel dimensions to generate an intermediate visual word sequence containing multi-granularity information. The sequence can be expressed as The process can be summarized as follows:
[0011]
[0012] Furthermore, in step 2, the purpose of spatial word fusion (STF) is to reduce the spatial redundancy between visual word elements, thereby shortening the length of the sequence input to the large language model. In this stage, the intermediate visual word sequence generated in the previous step (or directly the original visual word sequence from the visual encoder) is spatially fused. The specific operation is to perform spatial fusion on the input visual word sequence, such as It has shape H1×W1×C1 and applies a convolution operation conv with kernel size k×k and stride also k×k. k×k And followed by an activation function. This operation fuses every k×k spatially adjacent word units into a new, more compact word unit. For example, when k=2, the 2×2 word unit neighborhood is fused into one word unit. After fusion, the initial fused word unit sequence is obtained. Its shape is H2×W2×C2, where H2=H1 / k and W2=W1 / k. If this convolution operation is considered as a learnable concatenation operation, the output channel C2 can be set to k 2 C1. In addition, in order to more flexibly control the number and dimension of output tokens, Reshape the tensor into E tokens, each of which has the shape Among them E <k 2 .
[0013] Furthermore, in step 3, the alignment of the compact visual word unit with the language model is achieved by a further projection operation. For example, one or more convolutional layers with a kernel size of 1×1 followed by an activation function are used to align the dimensions of the compact visual word unit with the text embedding space of the large-scale language model, so that its dimensions are consistent with the text word unit dimension C of the language model. LLM (corresponding to C3 in the paper) consistent. Finally, the aligned compact visual word sequence is obtained The total number of words is reduced from the original H1·W1 to
[0014] Furthermore, in step 4, during model inference, the final aligned compact visual word obtained in step 3 is The text words corresponding to the user instructions are input into the large-scale language model h φ To generate the corresponding text response. The model optimization adopts a two-stage training strategy. The first stage is feature alignment pre-training, in which only the learnable parameters θ of the MBTF and STF modules and the alignment projection layer (collectively referred to as the projector g θ ), while the visual encoder and large-scale language model h φ The parameters of are kept fixed. The second stage is instruction fine-tuning, in which the projector g is updated. θ The parameters θ and the large-scale language model h φ All or part of the parameters φ of φ can be tuned (e.g., through full parameter fine-tuning or parameter-efficient fine-tuning methods such as LoRA), while the parameters of the visual encoder remain fixed. The training of both stages aims to maximize the generated target response X response The probability is calculated according to the following formula:
[0015]
[0016] where X v is the original visual information corresponding to the input image, X instruct For user instructions, X response,<i is the generated partial response, x i is the currently predicted response token.
[0017] Beneficial effects
[0018] The proposed efficient image analysis and question-answering system for intelligent customer service brings significant commercial value to intelligent customer service applications, successfully resolving the difficult balance between "response speed" and "answer accuracy" in online shopping scenarios. Its most direct benefits are reflected in two aspects: user experience and operating costs. First, through exponential efficiency improvements, it ensures that customers receive immediate feedback after uploading images and asking questions, avoiding customer churn caused by waiting and directly improving purchase conversion rates. Second, the system's efficient operation significantly reduces reliance on expensive computing resources, enabling e-commerce platforms to deploy high-quality visual question-answering customer service on a large scale at a lower cost, achieving cost reduction and efficiency improvement. Furthermore, through a unique information refinement mechanism, the system ensures high-precision recognition of product details while maintaining high-speed response, ensuring accurate answers. Its modular design also allows it to be easily integrated into existing e-commerce customer service platforms as an "efficiency accelerator," significantly reducing the threshold and cost of technology upgrades. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is an overall schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0020] The following is a detailed description of the embodiments of the present invention with reference to the accompanying drawings and specific experimental settings. The specific implementation of the present invention is described using an application based on the LLaVA-1.5-7B model as an example:
[0021] Step 1: Select and configure model components. First, select an appropriate visual encoder and large language model as the foundation. In this example, the visual encoder uses CLIP ViT-L / 14, with an input image resolution of 336×336, and its parameters remain frozen throughout training and inference. The large language model (LLM) uses Vicuna-1.5-7B as its backbone.
[0022] Secondly, configure the multi-block word fusion MBTF module. This module uniformly selects the output of (M=8) blocks from the visual encoder (ViT-L / 14 has a total of 24 Transformer blocks) as input features, for example, the output of the 3rd, 6th, 9th, 12th, 15th, 18th, 21st, and 24th blocks. After reshaping the output visual word of each selected block (dimension is 1024) into a feature map, these 8 sets of feature maps are spliced together along the channel dimension. Then, two consecutive 1×1 convolutional layers are used for fusion, and each convolutional layer is followed by a GeLU activation function. The first convolutional layer reduces the number of channels to, for example, 4096, and the second convolutional layer further reduces it to 1024, which is consistent with the original visual word dimension, and we get
[0023] Next, configure the spatial word fusion STF module. Its spatial dimension is, for example, 24×24 and the channel dimension is 1024 as input. The default setting is fusion kernel size k=2, and the target fusion word number E=1. A 2×2 convolution (with a step size of 2) is used to Processing is performed to fuse the 2×2×1024 features into 1×1×C2 features. If C2 is designed to be k 2 C1 = 4 1024 = 4096, then we can directly get the fused word that matches the LLM text embedding dimension (for example, the text embedding dimension of Vicuna-1.5-7B is 4096). Then, two 1×1 convolutional layers (followed by GeLU) are used to perform dimension alignment and feature transformation, and the final output is a compact visual word sequence. The dimension of each word unit is 4096. At this point, the number of visual word units is reduced to the original (1 / k) 2 =1 / 4.
[0024] Step 2: Model training. The model training follows the two-stage training scheme of LLaVA-1.5.
[0025] The first stage is the feature alignment pre-training stage. In this stage, the dataset used is a filtered CC-595K subset. The AdamW optimizer is used without weight decay, and the initial learning rate is set to 1×10 -3 , the batch size is 256. The learning rate adopts the cosine annealing strategy with a warm-up ratio of 3%. The training is performed for 1 epoch. At this stage, only the parameters θ of the MBTF, STF modules and the final aligned projection layer are updated, while the LLM (h φ ) and the parameters of the visual encoder remain frozen.
[0026] The second stage is the instruction fine-tuning stage. In this stage, the dataset used is the LLaVA-Instruct-158K dataset. The initial learning rate is set to 2×10 -5, the batch size is 128, and the rest of the optimizers and learning rate scheduling strategies are the same as those in the pre-training stage. Training also proceeds for one epoch. In this stage, the parameters θ of the MBTF, STF, and the aligned projection layer are updated, as well as the LLM (h φ ), the visual encoder parameters remain frozen. To conserve GPU memory, DeepSpeed and gradient checkpointing can be used during training without parameter offloading, and mixed-precision training with bfloat16 and TensorFloat32 precision can be enabled.
[0027] Step 3: Performance Evaluation and Results Analysis: The model configured and trained according to the above steps is evaluated on eight popular vision-language benchmarks (GQA, SQA, VQAv2, VisWiz, TextVQA, POPE, MMBench, and MMBench-CN).
[0028] Evaluation results show that compared to the baseline LLaVA-1.5 (full visual tokens), our method achieves comparable or even slightly higher average scores (e.g., 0.3% to 0.8% higher average scores) when using only 25% of visual tokens. This demonstrates that our method can effectively maintain or even improve the model's multimodal reasoning capabilities while significantly reducing visual tokens. Compared with other efficient LLaVA models (such as PruMerge+, FastV, and YOPO) at a similar computational cost (approximately 1.9 TFLOPs), our method achieves the best average performance.
[0029] Ablation experiments further verified the effectiveness of the MBTF and STF modules. Using only the MBTF module (without reducing the number of tokens) has improved performance compared to the baseline, illustrating the importance of multi-granularity features. Using only the STF module (with the number of tokens reduced to 25%) can also achieve performance comparable to the baseline, proving that there is significant spatial redundancy in visual tokens. The combination of the two achieves the best performance balance while significantly reducing the amount of computation. Studies on the fusion kernel size k and the number of fused tokens E of the STF module show that the configuration of (k=2, E=1), that is, 2×2 neighborhoods are fused into 1 token, is the best. Too large a fusion kernel or too few fused tokens may lead to information loss and affect performance. Compared with different fusion strategies (such as average pooling AvgPool and simple concatenation TokenConcat), the learnable STF module proposed in this invention performs better and can adaptively fuse visual information.
[0030] Through the above specific implementation methods, the present invention proposes an efficient image analysis and question-answering system for intelligent customer service, which transforms complex AI models into a practical solution with faster response and easier deployment by combining advanced visual understanding technology with an optimized operating architecture.
Claims
1. An efficient image analysis and question-answering system for intelligent customer service, comprising the following steps: Step 1: extract multiple sets of visual feature words from multiple pre-selected layers of the visual encoder, and perform multi-block word fusion on the multiple sets of visual feature words to fuse them into an intermediate visual word sequence containing multi-granularity information; Step 2: Perform spatial word fusion on the intermediate visual word sequence generated in step 1, or directly on the original visual word sequence from the visual encoder, to aggregate multiple spatially adjacent visual words into a smaller number of compact visual words; Step 3: Align the compact visual word generated in step 2 to the text embedding space of the large-scale language model through a projection operation to obtain the final aligned compact visual word, which is recorded as Step 4: The final aligned compact visual word obtained in step 3 is The text words corresponding to the user instructions are input into the large-scale language model to generate a response.
2. The method according to claim 1, characterized in that The multi-block word fusion in step 1 specifically includes: (1) From the M pre-selected layers of the visual encoder Extract respective visual words (2) Extract each visual word Reshape into a feature map with predetermined spatial and channel dimensions; (3) splicing the reshaped feature maps along their channel dimension; (4) Using at least one convolution operation with a kernel size of 1×1 followed by an activation function, such as GeLU, to process the concatenated features to fuse features and adjust channel dimensions to generate the intermediate visual word sequence, which can be expressed as The process can be summarized as 3. The method according to claim 1, characterized in that The spatial word unit fusion in step 2 specifically includes: (1) For the input visual word sequence, for example It has shape H1×W1×C1 and applies a convolution operation conv with kernel size k×k and stride also k×k. k×k And then an activation function such as GeLU is used to fuse each k×k adjacent word into one or more new words to obtain a preliminary fused word sequence Its shape is (H2×W2×C2), where (H2=H1 / k), (W2=W1 / k); (2) Optionally, Reshape the tensor to adjust the number of output tokens to E, and the shape of each token is where E is less than or equal to k 2 .
4. The method according to claim 1, wherein The projection operation in step 3 specifically includes: For the compact visual word obtained in step 3 (e.g. or its reshaped form) applies at least one convolution operation with a kernel size of 1×1 followed by an activation function to resize its feature dimension to the text embedding dimension C of the large-scale language model LLM (corresponding to C3 in the paper), thus obtaining the final aligned compact visual word The total number is 5. The method according to claim 1, wherein The method is optimized through a two-stage training process, which includes: The first stage is feature alignment pre-training: the parameters θ are learned by simply updating the modules that perform the fusion and projection operations described in steps 1, 2, and 3, collectively referred to as the projector g. θ , while the visual encoder and large-scale language model h φ The parameters of remain fixed; The second stage is instruction fine-tuning: Update the projector g θ The parameters θ and the large-scale language model h φ The parameters φ of the visual encoder are kept fixed; and the training of the two stages aims to maximize the generated target response X response The probability is calculated according to the following formula: where X v is the original visual information corresponding to the input image, X instruct For user instructions, X response,<i is the generated partial response, x i is the currently predicted response token.
Citation Information
Cited By
An industrial safety operation and maintenance large model-oriented question and answer method and system
CN122549601B