Multi-modal model-based training method of retail store front photo comparison model

Through the training method of retail store door-head photo comparison model based on multimodal model, the weight is adjusted using structured annotation data, and the problem of low accuracy in door-head photo comparison in the existing technology is solved, achieving more efficient and accurate door-head photo comparison.

CN120147790APending Publication Date: 2025-06-13BEIJING PRISM INTELLIGENT TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510629176.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When comparing photos of retail stores, the prior art has the problem of low accuracy in comparison due to the photos containing pedestrians, vehicles and other obstacles, and the preprocessing steps of the multimodal model are cumbersome, which increases costs.

Method used

By providing a training method for retail store door head photo comparison model based on multimodal model, the door head area, fixed area around the door head and obstacle areas are adjusted by using structured annotation data, and the trained model is fine-tuned to focus on the door head area and ignore obstacles.

Benefits of technology

The preprocessing steps for photo comparison are reduced, cost is reduced, and the accuracy of comparison is improved, allowing the model to more accurately identify and compare retail store front photos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147790A_ABST
    Figure CN120147790A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a retail store front photo comparison model based on a multi-modal model, and relates to the technical field of retail industry image processing. The method comprises the following steps: performing fine tuning training on a preset multi-modal model, and inserting a first weight corresponding to the position and content of a door head region, a second weight corresponding to the position and content of a fixed region around the door head, and a third weight corresponding to the position and content of an obstacle region during fine tuning training to obtain a retail store door head picture comparison model after fine tuning training; according to the method, the retail shop door head picture comparison model after fine tuning training can focus attention on the position and content of a door head area and the position and content of a fixed area around the door head, and interferents such as pedestrians, vehicles or other objects are neglected, so that the preprocessing steps can be reduced, the door head picture comparison cost can be reduced, and the comparison accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of image processing in the retail industry, and particularly to a training method for a retail store facade photo comparison model based on a multi-modal model. Background Art

[0002] In retail industry management, in order to ensure that salespersons actually arrive at the retail stores they are responsible for to carry out their work, photo check-in is generally used for supervision and confirmation. The specific approach is to compare a standard retail store facade photo with the retail store facade photo taken by the salesperson to determine whether the salesperson has truly arrived at the designated store.

[0003] In the prior art, when comparing a standard retail store facade photo with the retail store facade photo taken by the salesperson, it is mainly based on a pixel-based comparison method, that is, comparing the pixel value differences of the two photos pixel by pixel; and when the retail store facade photo taken by the salesperson contains objects that do not exist in the standard retail store facade photo (such as pedestrians and vehicles that can move), there will be interference and the accuracy of the comparison is not high.

[0004] In addition, for multi-modal model comparison, compared with the pixel-based comparison method, the multi-modal model can process more complex and variable photos and has the ability to analyze and understand the meaning of the photos; however, if the multi-modal model is directly used, the facade photos need to be preprocessed first, including cropping, rotation, contrast adjustment, etc., which highly depends on various preprocessing steps. For example: a check-in photo contains the retail store facade, the storefront, and there is a car in front of the storefront, while the standard retail store facade photo only contains the facade and some content near the facade. If directly compared, the similarity will be very low. Therefore, before the photo is input into the multi-modal model, the photo needs to be cropped to ensure that a large area of the photo is the facade and the storefront and the car in front of the storefront are removed, and then the photo can be described and input. Therefore, how to optimize the multi-modal model to reduce the preprocessing steps, lower the cost of retail store facade photo comparison, and improve the comparison accuracy has become a technical problem to be urgently solved. Summary of the Invention

[0005] In view of the above problems, this application is proposed to provide a training method, device, equipment and medium for a retail store facade photo comparison model based on a multi-modal model to overcome or at least partially solve the above problems. The technical solutions are as follows: In a first aspect, a training method for a retail store facade photo comparison model based on a multi-modal model is provided. The method enables the fine-tuned retail store facade photo comparison model to focus on the facade area position and content, and the fixed area position and content around the facade. The method includes: Obtain standard retail store facade photos, and obtain data for structured annotation of the facade area position and content, the fixed area position and content around the facade of the standard retail store facade photos; Collect retail store facade photos in multiple scenarios, and obtain data for structured annotation of the facade area position and content, the fixed area position and content around the facade, and the obstacle area position and content of the retail store facade photos in the multiple scenarios; Input the standard retail store facade photos, the data for structured annotation of the facade area position and content, the fixed area position and content around the facade of the standard retail store facade photos, the retail store facade photos in the multiple scenarios, and the data for structured annotation of the facade area position and content, the fixed area position and content around the facade, and the obstacle area position and content of the retail store facade photos in the multiple scenarios into a preset multimodal model, perform fine-tuning training on the preset multimodal model, and insert a first weight corresponding to the facade area position and content, a second weight corresponding to the fixed area position and content around the facade, and a third weight corresponding to the obstacle area position and content during the fine-tuning training to obtain a fine-tuned retail store facade photo comparison model; wherein, the third weight is less than the first weight, and the third weight is less than the second weight.

[0006] In a possible implementation manner, after obtaining the fine-tuned retail store facade photo comparison model, the method further includes: When receiving an actual retail store facade photo taken by a user, obtain the current standard retail store facade photo of the retail store where the user is located, input the actual retail store facade photo and the current standard retail store facade photo into the fine-tuned retail store facade photo comparison model, and output a structured feature vector of the actual retail store facade photo and a structured vector of the current standard retail store facade photo, wherein the structured feature vector of the actual retail store facade photo includes the facade area content and the fixed area content around the facade of the actual retail store facade photo, and the structured vector of the current standard retail store facade photo includes the facade area content and the fixed area content around the facade of the current standard retail store facade photo; Calculate the similarity between the actual retail store facade photo and the current standard retail store facade photo according to the structured feature vector of the actual retail store facade photo and the structured vector of the current standard retail store facade photo; Determine whether the actual retail store facade photo and the current standard retail store facade photo are facade photos of the same retail store according to the similarity between the actual retail store facade photo and the current standard retail store facade photo.

[0007] In a possible implementation manner, determining whether the actual retail store front door photo and the current standard retail store front door photo are the front door photos of the same retail store according to the similarity between the actual retail store front door photo and the current standard retail store front door photo includes: If the similarity between the actual retail store front door photo and the current standard retail store front door photo is greater than a preset threshold, it is determined that the actual retail store front door photo and the current standard retail store front door photo are the front door photos of the same retail store; If the similarity between the actual retail store front door photo and the current standard retail store front door photo is less than or equal to the preset threshold, it is determined that the actual retail store front door photo and the current standard retail store front door photo are not the front door photos of the same retail store.

[0008] In a possible implementation manner, collecting retail store front door photos in multiple situations, including one or more of the following: Collecting photos of the same retail store from different angles and at different distances; Collecting photos of the same retail store with vehicles in front of it; Collecting photos of the same retail store with pedestrians in front of it; Collecting photos of the same retail store with an unobstructed front; Collecting photos in which the proportion of the retail store front door in the photo is within a set proportion range.

[0009] In a possible implementation manner, when collecting photos of the same retail store from different angles and at different distances, the data obtained by performing structured annotation on the retail store front door photos in multiple situations for the position and content of the front door area, the position and content of the fixed area around the front door, and the position and content of the obstacle area also includes descriptions of the front door angle and distance; When collecting photos of the same retail store with vehicles in front of it, the data obtained by performing structured annotation on the retail store front door photos in multiple situations for the position and content of the front door area, the position and content of the fixed area around the front door, and the position and content of the obstacle area also includes a description of the position of the vehicle relative to the front door; When collecting photos of the same retail store with pedestrians in front of it, the data obtained by performing structured annotation on the retail store front door photos in multiple situations for the position and content of the front door area, the position and content of the fixed area around the front door, and the position and content of the obstacle area also includes a description of the position of the pedestrian relative to the front door; When collecting photos of the same retail store with an unobstructed front, the data obtained by performing structured annotation on the retail store front door photos in multiple situations for the position and content of the front door area, the position and content of the fixed area around the front door, and the position and content of the obstacle area also includes a description of the unobstructed front; When collecting photos with the proportion of the storefront in the same retail store within the set proportion range, the data obtained by structuring and annotating the storefront area position and content, the fixed area position and content around the storefront, and the obstacle area position and content of the storefront photos in the above-mentioned multiple situations also includes a description of the proportion of the storefront in the photo.

[0010] In a possible implementation manner, the preset multimodal model is fine-tuned and trained, and when fine-tuning and training, a first weight corresponding to the storefront area position and content, a second weight corresponding to the fixed area position and content around the storefront, and a third weight corresponding to the obstacle area position and content are inserted to obtain a fine-tuned and trained storefront photo comparison model for retail stores, including: Load the Low-Rank Adaptation (LoRA) module to fine-tune and train the preset multimodal model. When fine-tuning and training, insert a first weight corresponding to the storefront area position and content, a second weight corresponding to the fixed area position and content around the storefront, and a third weight corresponding to the obstacle area position and content, update the parameters of the LoRA module, and freeze the main parameters of the preset multimodal model to obtain a fine-tuned and trained storefront photo comparison model for retail stores.

[0011] In a possible implementation manner, in the data obtained by structuring and annotating the storefront area position and content, the fixed area position and content around the storefront of the standard storefront photo of the retail store, the storefront area position and the fixed area position around the storefront are respectively represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the standard storefront photo of the retail store; In the data obtained by structuring and annotating the storefront area position and content, the fixed area position and content around the storefront, and the obstacle area position and content of the storefront photos in the above-mentioned multiple situations, the storefront area position, the fixed area position around the storefront, and the obstacle area position are respectively represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the storefront photos in the above-mentioned multiple situations.

[0012] In a second aspect, a training device for a storefront photo comparison model for retail stores based on a multimodal model is provided. The device enables the fine-tuned and trained storefront photo comparison model for retail stores to focus on the storefront area position and content, the fixed area position and content around the storefront. The device includes: A first acquisition unit, configured to acquire a standard storefront photo of a retail store and acquire data obtained by structuring and annotating the storefront area position and content, the fixed area position and content around the storefront of the standard storefront photo of the retail store; A second acquisition unit, configured to collect retail store facade photos in multiple scenarios, and acquire data of structured annotations for the positions and contents of the facade areas, the positions and contents of the fixed areas around the facades, and the positions and contents of the obstacle areas in the retail store facade photos in the multiple scenarios; A fine-tuning training unit, configured to input the standard retail store facade photos, the data of structured annotations for the positions and contents of the facade areas, the positions and contents of the fixed areas around the facades in the standard retail store facade photos, the retail store facade photos in the multiple scenarios, and the data of structured annotations for the positions and contents of the facade areas, the positions and contents of the fixed areas around the facades, and the positions and contents of the obstacle areas in the retail store facade photos in the multiple scenarios into a preset multimodal model, perform fine-tuning training on the preset multimodal model, and insert a first weight corresponding to the position and content of the facade area, a second weight corresponding to the position and content of the fixed area around the facade, and a third weight corresponding to the position and content of the obstacle area during fine-tuning training, to obtain a fine-tuned retail store facade photo comparison model; wherein, the third weight is less than the first weight, and the third weight is less than the second weight.

[0013] In a possible implementation manner, the apparatus further includes a comparison unit, configured to: After obtaining the fine-tuned retail store facade photo comparison model, when receiving an actual retail store facade photo taken by a user, acquire the current standard retail store facade photo of the retail store where the user is located, input the actual retail store facade photo and the current standard retail store facade photo into the fine-tuned retail store facade photo comparison model, and output a structured feature vector of the actual retail store facade photo and a structured vector of the current standard retail store facade photo, wherein, the structured feature vector of the actual retail store facade photo includes the content of the facade area and the content of the fixed area around the facade of the actual retail store facade photo, and the structured vector of the current standard retail store facade photo includes the content of the facade area and the content of the fixed area around the facade of the current standard retail store facade photo; Calculate the similarity between the actual retail store facade photo and the current standard retail store facade photo according to the structured feature vector of the actual retail store facade photo and the structured vector of the current standard retail store facade photo; Determine whether the actual retail store facade photo and the current standard retail store facade photo are facade photos of the same retail store according to the similarity between the actual retail store facade photo and the current standard retail store facade photo.

[0014] In a possible implementation manner, the comparison unit is further configured to: If the similarity between the actual retail store front door photo and the current standard retail store front door photo is greater than the preset threshold, it is determined that the actual retail store front door photo and the current standard retail store front door photo are the front door photos of the same retail store; If the similarity between the actual retail store front door photo and the current standard retail store front door photo is less than or equal to the preset threshold, it is determined that the actual retail store front door photo and the current standard retail store front door photo are not the front door photos of the same retail store.

[0015] In a possible implementation, the second acquisition unit is further configured to: Collect retail store front door photos in multiple scenarios, including one or more of the following: Collect photos of the same retail store from different angles and at different distances; Collect photos of the same retail store with vehicles in front of it; Collect photos of the same retail store with pedestrians in front of it; Collect photos of the same retail store with no obstruction in front of it; Collect photos of the same retail store where the proportion of the store front door in the photo is within a set proportion range.

[0016] In a possible implementation, when collecting photos of the same retail store from different angles and at different distances, the data obtained by performing structured annotation on the retail store front door photos in multiple scenarios for the position and content of the store front door area, the position and content of the fixed area around the store front door, and the position and content of the obstacle area further includes descriptions of the store front door angle and distance; When collecting photos of the same retail store with vehicles in front of it, the data obtained by performing structured annotation on the retail store front door photos in multiple scenarios for the position and content of the store front door area, the position and content of the fixed area around the store front door, and the position and content of the obstacle area further includes a description of the position of the vehicle relative to the store front door; When collecting photos of the same retail store with pedestrians in front of it, the data obtained by performing structured annotation on the retail store front door photos in multiple scenarios for the position and content of the store front door area, the position and content of the fixed area around the store front door, and the position and content of the obstacle area further includes a description of the position of the pedestrian relative to the store front door; When collecting photos of the same retail store with no obstruction in front of it, the data obtained by performing structured annotation on the retail store front door photos in multiple scenarios for the position and content of the store front door area, the position and content of the fixed area around the store front door, and the position and content of the obstacle area further includes a description of no obstruction in front of the store; When collecting photos with the proportion of the storefront of the same retail store in the set proportion range, the data obtained by structuring and annotating the storefront photos of the retail store in the above-mentioned multiple situations for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of the proportion of the storefront in the photo.

[0017] In a possible implementation manner, the fine-tuning training unit is further configured to: Load a Low-Rank Adaptation (LoRA) module to fine-tune and train the preset multi-modal model. During the fine-tuning training, insert the first weight corresponding to the position and content of the storefront area, the second weight corresponding to the position and content of the fixed area around the storefront, and the third weight corresponding to the position and content of the obstacle area, update the parameters of the LoRA module, and freeze the main parameters of the preset multi-modal model to obtain a fine-tuned and trained comparison model for retail storefront photos.

[0018] In a possible implementation manner, in the data obtained by structuring and annotating the position and content of the storefront area, the position and content of the fixed area around the storefront of the standard retail storefront photo, the position of the storefront area and the position of the fixed area around the storefront are respectively represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the standard retail storefront photo; In the data obtained by structuring and annotating the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area of the storefront photos of the retail store in the above-mentioned multiple situations, the position of the storefront area, the position of the fixed area around the storefront, and the position of the obstacle area are respectively represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the storefront photos of the retail store in the above-mentioned multiple situations.

[0019] In a third aspect, an electronic device is provided. The electronic device includes a processor and a memory. Among them, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the training method of the comparison model for retail storefront photos based on the multi-modal model described in any one of the above.

[0020] In a fourth aspect, a storage medium is provided. The storage medium stores a computer program. Among them, the computer program is configured to execute the training method of the comparison model for retail storefront photos based on the multi-modal model described in any one of the above when running.

[0021] With the above technical solutions, the training method, device, equipment and medium of the retail store facade photo comparison model based on the multi-modal model provided by the embodiments of the present application enable the fine-tuned retail store facade photo comparison model to focus on the facade area position and content, the fixed area position and content around the facade, and ignore the interferences such as pedestrians, vehicles or other objects, so as to reduce the preprocessing steps, reduce the cost of facade photo comparison, and improve the comparison accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments of the present application will be briefly introduced below.

[0023] Figure 1 Shows a flowchart of the training method of the retail store facade photo comparison model based on the multi-modal model provided by the embodiments of the present application; Figure 2 Shows a structural diagram of the training device of the retail store facade photo comparison model based on the multi-modal model provided by the embodiments of the present application; Figure 3 Shows a structural diagram of the training device of the retail store facade photo comparison model based on the multi-modal model provided by another embodiment of the present application; Figure 4 Shows a structural diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] The exemplary embodiments of the present application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "including" and its variants should be interpreted as an open term meaning "including but not limited to".

[0026] To solve the above technical problems, an embodiment of the present application provides a training method for a retail store facade photo comparison model based on a multimodal model, which enables the fine-tuned and trained retail store facade photo comparison model to focus on the position and content of the facade area, and the position and content of the fixed area around the facade. As Figure 1 shown, the training method for the retail store facade photo comparison model based on the multimodal model may include the following steps S101 to S103: Step S101, obtain a standard retail store facade photo, and obtain data for structurally annotating the position and content of the facade area, and the position and content of the fixed area around the facade of the standard retail store facade photo.

[0027] In this step, the facade area refers to the facade part of the retail store. For example, for a retail store facade photo, the facade area is a red cuboid background, and the content is the white font xx Convenience Store and the white logo (mark).

[0028] The fixed area around the facade refers to the part that does not change or move around the facade of the retail store. For example, for a retail store facade photo, there is a white air conditioner fixed above the facade.

[0029] Step S102, collect retail store facade photos in various situations, and obtain data for structurally annotating the position and content of the facade area, the position and content of the fixed area around the facade, and the position and content of the obstacle area of the retail store facade photos in various situations.

[0030] In this step, the facade area and the fixed area around the facade can be referred to the previous introduction. The obstacle area refers to the obstacle part at the entrance of the retail store, which will change and move. For example, for a retail store facade photo, there are pedestrians or cars at the facade.

[0031] It should be noted that the example introductions of the facade area, the fixed area around the facade, and the obstacle area are only illustrative and do not limit this embodiment.

[0032] Step S103: Input the standard retail store facade photos, the data of the structural annotation of the facade area position and content, the fixed area position and content around the facade of the standard retail store facade photos, the retail store facade photos in multiple scenarios, and the data of the structural annotation of the facade area position and content, the fixed area position and content around the facade, and the obstacle area position and content of the retail store facade photos in multiple scenarios into a preset multimodal model, fine-tune and train the preset multimodal model, and insert the first weight corresponding to the facade area position and content, the second weight corresponding to the fixed area position and content around the facade, and the third weight corresponding to the obstacle area position and content during the fine-tuning training to obtain a fine-tuned and trained retail store facade photo comparison model; wherein, the third weight is less than the first weight, and the third weight is less than the second weight.

[0033] In this step, the preset multimodal model is open-source. It is an artificial intelligence model that can process and integrate multiple data types (such as text, images, audio, video, etc.). Its core lies in improving the understanding and generation ability of complex tasks through cross-modal information fusion.

[0034] The third weight being less than the first weight and the second weight enables the fine-tuned and trained retail store facade photo comparison model to focus on the facade area position and content, the fixed area position and content around the facade, and ignore the obstacle area position and content, that is, ignore the interference objects such as pedestrians, vehicles, or other objects. Here, the third weight can be set to a very small value, such as set to 0 or 0.0001, etc. In a specific embodiment, the first weight, the second weight, and the third weight can be set to 0.8, 0.2, and 0 respectively. It should be noted that the examples here are only illustrative and do not limit this embodiment.

[0035] The training method of the retail store facade photo comparison model based on the multimodal model in this embodiment enables the fine-tuned and trained retail store facade photo comparison model to focus on the facade area position and content, the fixed area position and content around the facade, and ignore the interference objects such as pedestrians, vehicles, or other objects, thereby reducing the preprocessing steps, lowering the cost of facade photo comparison, and improving the comparison accuracy.

[0036] A possible implementation manner is provided in the embodiment of the present application. After step S103 obtains the fine-tuned and trained retail store facade photo comparison model, the following steps A1 to A3 may further be included: Step A1, when receiving the actual retail store front photo taken by the user, obtain the current standard retail store front photo of the retail store where the user is located. Input the actual retail store front photo and the current standard retail store front photo into the fine-tuned retail store front photo comparison model. Output the structured feature vector of the actual retail store front photo and the structured vector of the current standard retail store front photo. Among them, the structured feature vector of the actual retail store front photo includes the content of the store front area and the content of the fixed area around the store front of the actual retail store front photo. The structured vector of the current standard retail store front photo includes the content of the store front area and the content of the fixed area around the store front of the current standard retail store front photo; Step A2, calculate the similarity between the actual retail store front photo and the current standard retail store front photo according to the structured feature vector of the actual retail store front photo and the structured vector of the current standard retail store front photo; Step A3, determine whether the actual retail store front photo and the current standard retail store front photo are the store front photos of the same retail store according to the similarity between the actual retail store front photo and the current standard retail store front photo.

[0037] In the fine-tuned retail store front photo comparison model of this embodiment, when comparing retail store front photos, the store front area is mainly considered, such as the first weight is 0.8; the fixed area around the store front in the photo is secondarily considered, such as the second weight is 0.2. For example, there is an actual retail store front photo with a car in the middle, a store front above, an air conditioner above the store front, doors and windows below the store front, and some leaves beside the store front; in the conventional way, it is necessary to crop the car in the photo or perform other image operations, and then input it into the multi-modal model; in this embodiment, the actual retail store front photo is directly input into the fine-tuned retail store front photo comparison model, allowing the model itself to ignore the interference objects, such as pedestrians, vehicles or other objects, mainly focusing on the store front, and then considering the air conditioner, doors and windows, leaves around the store front again, reducing the preprocessing steps, reducing the cost of store front photo comparison, and improving the comparison accuracy.

[0038] A possible implementation manner is provided in the embodiment of the present application. In the above step A3, according to the similarity between the actual retail store front photo and the current standard retail store front photo, it is determined whether the actual retail store front photo and the current standard retail store front photo are the store front photos of the same retail store. Specifically, it may include the following steps A3-1 and A3-2: Step A3-1, if the similarity between the actual retail store front photo and the current standard retail store front photo is greater than the preset threshold, it is determined that the actual retail store front photo and the current standard retail store front photo are the store front photos of the same retail store; Step A3-2: If the similarity between the actual retail store facade photo and the current standard retail store facade photo is less than or equal to the preset threshold, it is determined that the actual retail store facade photo and the current standard retail store facade photo are not the facade photos of the same retail store.

[0039] In this step, the preset threshold can be set according to actual needs. For example, the preset threshold can be set to 80% or 85%, etc., and this embodiment does not limit this. When it is determined that the actual retail store facade photo and the current standard retail store facade photo are the facade photos of the same retail store, a first prompt message indicating that the actual retail store facade photo and the current standard retail store facade photo are the facade photos of the same retail store can be generated; when it is determined that the actual retail store facade photo and the current standard retail store facade photo are not the facade photos of the same retail store, a second prompt message indicating that the actual retail store facade photo and the current standard retail store facade photo are not the facade photos of the same retail store can be generated.

[0040] In an actual application scenario, there are 3 retail stores of the "Delicious and Refreshing" chain of beverages on a certain pedestrian street, and the air conditioners, doors and windows, leaves, etc. around the facades of these 3 retail stores are different; adopting the solution of this embodiment, mainly focusing on the facade and then considering the air conditioners, doors and windows, leaves, etc. around the facade again, the retail stores of the "Delicious and Refreshing" chain of beverages can be accurately identified, and these 3 retail stores can be accurately distinguished.

[0041] In an embodiment of the present application, a possible implementation manner is provided. The above step S102 of collecting retail store facade photos in multiple situations may include one or more of the following: Collecting photos of the same retail store from different angles and at different distances; Collecting photos of the same retail store with vehicles in front of it; Collecting photos of the same retail store with pedestrians in front of it; Collecting photos of the same retail store with no obstruction in front of it; Collecting photos of the same retail store where the proportion of the facade in the photo is within a set proportion range.

[0042] This embodiment collects retail store facade photos in multiple situations, which may include photos of the same retail store from different angles and at different distances, photos of the same retail store with vehicles in front of it, photos of the same retail store with pedestrians in front of it, photos of the same retail store with no obstruction in front of it, photos of the same retail store where the proportion of the facade in the photo is within a set proportion range, etc. In this way, it is closer to the real situation, strengthens the data proportion in complex scenarios (such as angles, obstructions, lighting, etc.), and ensures that the retail store facade photo comparison model after fine-tuning training can stably focus on the facade in the real environment of attendance check-in.

[0043] In an embodiment of the present application, a possible implementation is provided. When collecting photos of the same retail store from different angles and at different distances, the data obtained by performing structured annotation on the storefront photos of the retail store in multiple scenarios for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes descriptions of the storefront angle and the distance. When collecting photos of the same retail store with vehicles in front of it, the data obtained by performing structured annotation on the storefront photos of the retail store in multiple scenarios for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of the position of the vehicle relative to the storefront. When collecting photos of the same retail store with pedestrians in front of it, the data obtained by performing structured annotation on the storefront photos of the retail store in multiple scenarios for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of the position of the pedestrian relative to the storefront. When collecting photos of the same retail store with no obstruction in front of it, the data obtained by performing structured annotation on the storefront photos of the retail store in multiple scenarios for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of no obstruction in front of the store. When collecting photos in which the proportion of the storefront of the same retail store in the photo is within a set proportion range, the data obtained by performing structured annotation on the storefront photos of the retail store in multiple scenarios for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of the proportion of the storefront in the photo. The set proportion range here can be set according to actual needs. For example, the set proportion range can be the interval [1 / 4, 1 / 2]. It should be noted that the examples listed here are only illustrative and do not limit this embodiment.

[0044] In an embodiment of the present application, a possible implementation is provided. In the above step S103, the preset multimodal model is fine-tuned, and when fine-tuning, the first weight corresponding to the position and content of the storefront area, the second weight corresponding to the position and content of the fixed area around the storefront, and the third weight corresponding to the position and content of the obstacle area are inserted to obtain the fine-tuned retail storefront photo comparison model. Specifically, it may include the following steps S103-1: Step S103-1, load the Low-Rank Adaptation (LoRA) module to fine-tune the preset multimodal model. When fine-tuning, insert the first weight corresponding to the position and content of the storefront area, the second weight corresponding to the position and content of the fixed area around the storefront, and the third weight corresponding to the position and content of the obstacle area, update the parameters of the LoRA module, and freeze the main parameters of the preset multimodal model to obtain the fine-tuned retail storefront photo comparison model.

[0045] In this step, LoRA stands for Low Rank Adaptation, which is translated into Chinese as low-rank adaptation. Its core idea is to perform low-rank decomposition on a preset multimodal model, and only a small number of parameters need to be trained to achieve efficient fine-tuning. Specifically, by adding low-rank matrices (such as a and b) beside the weights of the preset multimodal model, the incremental changes of full-parameter fine-tuning can be simulated, rather than directly modifying the original parameters. This design can not only maintain the knowledge integrity of the preset multimodal model but also significantly reduce the demand for computing resources.

[0046] In addition, the fine-tuning parameters of LoRA, such as the learning rate and the number of iterations, can be adjusted to further optimize the model performance.

[0047] In an embodiment of the present application, a possible implementation is provided. In the data of the structural annotation of the position and content of the storefront area and the position and content of the fixed area around the storefront in the standard storefront photos of retail stores, the position of the storefront area and the position of the fixed area around the storefront are represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the standard storefront photos of retail stores. In the data of the structural annotation of the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area in the storefront photos of retail stores in multiple cases, the position of the storefront area, the position of the fixed area around the storefront, and the position of the obstacle area are represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the storefront photos of retail stores in multiple cases.

[0048] For example, the position of the storefront area is represented as rectangular coordinates [x1, y1, x2, y2], the position of the fixed area around the storefront is represented as rectangular coordinates [x3, y3, x4, y4], and the position of the obstacle area is represented as rectangular coordinates [x5, y5, x6, y6], and the coordinate values are the number of pixels starting from the upper left corner of the photo.

[0049] This embodiment can improve the accuracy and robustness of the recognition of the storefront and its surrounding areas. Through structural annotation and standardized coordinates, the model can better process storefront photos in various situations and is applicable to scenarios that require high-precision positioning and content analysis.

[0050] The above introduces Figure 1 Multiple implementation methods of each link of the illustrated embodiment. Next, the training method of the retail storefront photo comparison model based on the multimodal model of the embodiment of the present application will be further described through specific embodiments.

[0051] In this specific embodiment, the process of data preparation and fine-tuning the pre-set multi-modal model using LoRA to obtain the fine-tuned retail store facade photo comparison model will be introduced in detail. The solutions of directly using the pre-set multi-modal model, pre-processing plus using the pre-set multi-modal model, and the solution of using the fine-tuned retail store facade photo comparison model in this embodiment will also be analyzed and compared from multiple technical dimensions (such as core logic, attention distribution, data utilization, etc.) and evaluation dimensions (such as accuracy in normal scenarios, accuracy in occlusion scenarios, accuracy in angle change scenarios, human-like judgment fit, misjudgment types, etc.).

[0052] (1) Core concepts: Facade area: The facade part of a retail store. For example, in a photo of a retail store facade, the facade area is a red cuboid background with the content being the white text "XX Convenience Store" and the white logo (identification).

[0053] Fixed area around the facade: The part around the retail store facade that does not change or move. For example, in a photo of a retail store facade, there is a white air conditioner fixed above the facade.

[0054] Obstacle area: The obstacle part at the entrance of the retail store that changes and moves. For example, in a photo of a retail store facade, there are pedestrians or cars in front of the facade.

[0055] See Table 1, which shows the functions and annotation contents of the facade area, the fixed area around the facade, and the obstacle area.

[0056] Table 1

[0057] An example of the structured annotation format is as follows: { "image_path": "shop_001.jpg", "regions": {"type": "ROI-1", "coordinates": "[x1,y1,x2,y2]", "text": "XX Convenience Store red logo facade", "attributes": ["logo+text", "lighted characters"]}, {"type": "ROI-2", "coordinates": "[x3,y3,x4,y4]", "text": "Air conditioner, leaves above the facade, doors and windows below the facade"}, {"type": "ROI-3", "coordinates": "[x5,y5,x6,y6]", "text": "Black car in front of the store (ignored)"} } As shown in the above example, the photo is saved as shop_001.jpg, and the structured annotation data regions include the location and content of the storefront area, the location and content of the fixed area around the storefront, and the location and content of the obstacle area. It should be noted that the above example is only illustrative and does not limit this embodiment.

[0058] (2) This embodiment includes the following functions: 1) Data processing and enhancement module Basic annotation: Generate a text description containing core information for each storefront photo.

[0059] Region annotation: Use a preset annotation tool to mark the specific locations of the storefront area, the fixed area around the storefront, and the obstacle area in the photo.

[0060] Attribute content annotation: Extract the key attribute content tags of the storefront, such as "logo + text", "pure text", etc.

[0061] Geometric transformation: Randomly rotate the photo (such as ±15°), scale (0.8 - 1.2 times), and translate (such as ±10%) to simulate different shooting angles.

[0062] Image quality enhancement: Use a deblurring algorithm (such as non-local means filtering) to repair low-quality image photos and sharpen to highlight the text edges.

[0063] Semantic enhancement: Add random noise to the image photo and adjust the color temperature (such as ±2000K) to improve the robustness of the model to complex environments.

[0064] 2) LoRA configuration Rank configuration of the low-rank matrix: Rank value: Set r = 32 to balance the capture of storefront details and computational efficiency.

[0065] Target layer configuration: Such as the q_proj and v_proj layers.

[0066] Apply a weight mask to the feature vectors of the storefront area (ROI-1 area): In the b matrix of LoRA, set the row weight corresponding to the storefront area (ROI-1 area) to 0.8, the fixed area around the storefront (ROI-2 area) to 0.2, and the obstacle area (ROI-3 area) to 0.

[0067] 3) Training module, input the dataset into the model for training, and the steps are as follows: 3.1) Prepare standard data: Describe the standard retail storefront photos.

[0068] ​3.2) Collect the photos of the retail store facades under different conditions: Structurally annotate the location of the facade area, the location of the fixed areas around the facade, the location of the obstacle areas, the proportion of the picture occupied, etc.

[0069] Collect the photos of the same retail store from different angles and at different distances. The description mentions the distance of the facade, the area where the facade is located, and the fixed areas around the facade, such as "XXX Convenience Store, with a blue logo and a white-font facade. The facade is relatively far away, located in the upper left corner, there is an air conditioner and leaves above the facade, and there are doors and windows below the facade".

[0070] Collect the photos of the same retail store with vehicles in front of it. The description mentions the location of the vehicle relative to the facade, the area where the facade is located, and the fixed areas around the facade, such as "XXX Convenience Store, with a red logo and a white-font facade. There is a black car (to be ignored) below the facade, the facade is above the black car, and there is an air conditioner above the facade".

[0071] Collect the photos of the same retail store with pedestrians in front of it. The description mentions the location of the pedestrians relative to the facade, the area where the facade is located, and the fixed areas around the facade, such as "XXX Convenience Store, with a green logo and a white-font facade. There are three pedestrians (to be ignored) below the facade, the facade is directly above the pedestrians, and there is a green window above the facade".

[0072] Collect the photos of the same retail store with an unobstructed front. The description mentions that the facade is unobstructed and the fixed areas around the facade, such as "White logo and white-font facade, the facade is unobstructed, and the surrounding area is white".

[0073] Collect the photos of the same retail store where the facade occupies 1 / 4 - 1 / 2 of the photo. The description mentions the proportion of the facade in the photo and the fixed areas around the facade, such as "Blue logo and white-font facade, the facade occupies 1 / 4 of the photo, and there is a gray wall on the right side of the facade".

[0074] 3.3) Train the pre-set multi-modal model loaded with LoRA to ensure that during the training process, only the parameters of the LoRA module are updated, while the parameters of the pre-set multi-modal model remain frozen. Specifically, enhance the weight of the facade features through region-related semantics.

[0075] 3.4) Adjust the fine-tuning parameters of the low-rank adaptation LoRA, such as the learning rate, the number of iterations, etc., to further optimize the model performance.

[0076] 4) Normalize the photo feature vectors and calculate the similarity based on the normalized values. If the similarity is greater than 80%, it is determined that the photos are of the facades of the same retail store.

[0077] (3) The data collection method of this embodiment: 1) Core principle: Goal-driven: All data collection and annotation steps revolve around "priority of the storefront", with the design that the most refined annotation (ROI-1) has the highest data volume ratio.

[0078] Well-structured: Through three-level regional annotation, the core, secondary auxiliary, and irrelevant areas are explicitly distinguished, reducing the model's learning cost.

[0079] Close to reality: Strengthen the data ratio of complex scenarios (angles, occlusions, lighting) to ensure that the model can stably focus on the storefront in the real environment of attendance check-in.

[0080] 2) Define the scope of data collection 2.1) Main collection scenarios: 2.1.1) Storefront main body: Storefronts with clear store names and logos, such as the storefront of "XX Convenience Store", with brand logos and fonts.

[0081] 2.1.2) Complex scenarios (each type accounts for 10%-20%): Angle changes: Aerial photography, looking up, side shooting (tilted ±15°).

[0082] Occlusion situations: There are leaves, billboards, and sunshades around the storefront (occlusion area ≤ 30%).

[0083] Proportion of the storefront: Occupies 1 / 4 - 1 / 2 of the photo area (simulating far or near shooting).

[0084] 2.2) Auxiliary collection objects: Surrounding environment of the storefront: Air conditioners, doors and windows, leaves, walls, door frames (for 20% attention training).

[0085] 2.3) Excluded objects: Photos with the storefront completely occluded (occlusion > 50%), non-store photos (such as pure vehicle or pedestrian photos).

[0086] 3) Multi-source data collection 3.1) Collection channels: Physical shooting, historical check-ins.

[0087] 3.2) Collection tools: Hardware: Mobile phone (main camera ≥ 12 million pixels), tripod (fixed angle).

[0088] Software: Annotation tools, such as CVAT (Computer Vision Annotation Tool), which is an open-source and free annotation tool that supports various annotation types (such as bounding boxes, polygons, image segmentation, etc.) and is suitable for tasks such as object detection and semantic segmentation.

[0089] Referring to Table 2, it shows an analysis and comparison of the solutions of directly using a preset multi-modal model, the solution of preprocessing plus using a preset multi-modal model, and the solution of using the retail store facade photo comparison model after fine-tuning training in this embodiment from multiple technical dimensions (such as core logic, attention distribution, data utilization, etc.).

[0090] Table 2

[0091] Referring to Table 3, it shows an analysis and comparison of the solutions of directly using a preset multi-modal model, the solution of preprocessing plus using a preset multi-modal model, and the solution of using the retail store facade photo comparison model after fine-tuning training in this embodiment from multiple evaluation dimensions (such as normal scene accuracy, occluded scene accuracy, angle change accuracy, human-like judgment fit, misjudgment type, etc.).

[0092] Table 3

[0093] In this embodiment, the model is adjusted through fine-tuning training, so that when comparing retail store facade photos, the model can focus its attention on the facade and the fixed area near the facade, without considering other obstacles in the photo, enabling the trained and adjusted retail store facade photo comparison model to accurately identify and compare facades at different distances and angles, while reducing the comparison cost and simplifying the preprocessing process.

[0094] It should be noted that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. In practical applications, all the above possible implementation manners can be combined in any combination to form possible embodiments of the present application, which will not be elaborated herein one by one.

[0095] Based on the training method of the retail store facade photo comparison model based on a multi-modal model provided in the above embodiments, based on the same inventive concept, the embodiments of the present application also provide a training device for the retail store facade photo comparison model based on a multi-modal model.

[0096] Figure 2 It is a structural diagram of the training device for the retail store facade photo comparison model based on a multi-modal model provided in the embodiments of the present application. This device enables the retail store facade photo comparison model after fine-tuning training to focus its attention on the position and content of the facade area, and the position and content of the fixed area around the facade. AsFigure 2 As shown, the training device of the retail store facade photo comparison model based on the multi-modal model may specifically include a first acquisition unit 210, a second acquisition unit 220, and a fine-tuning training unit 230.

[0097] The first acquisition unit 210 is configured to acquire a standard retail store facade photo and acquire data for structurally annotating the position and content of the facade area, the position and content of the fixed area around the facade of the standard retail store facade photo. The second acquisition unit 220 is configured to collect retail store facade photos in multiple situations and acquire data for structurally annotating the position and content of the facade area, the position and content of the fixed area around the facade, and the position and content of the obstacle area of the retail store facade photos in the multiple situations. The fine-tuning training unit 230 is configured to input the standard retail store facade photo, the data for structurally annotating the position and content of the facade area, the position and content of the fixed area around the facade of the standard retail store facade photo, the retail store facade photos in the multiple situations, and the data for structurally annotating the position and content of the facade area, the position and content of the fixed area around the facade, and the position and content of the obstacle area of the retail store facade photos in the multiple situations into a preset multi-modal model, perform fine-tuning training on the preset multi-modal model, and insert a first weight corresponding to the position and content of the facade area, a second weight corresponding to the position and content of the fixed area around the facade, and a third weight corresponding to the position and content of the obstacle area during the fine-tuning training to obtain a fine-tuned retail store facade photo comparison model; wherein, the third weight is less than the first weight, and the third weight is less than the second weight.

[0098] In an embodiment of the present application, a possible implementation manner is provided. As Figure 3 shown, the device shown above Figure 2 may further include a comparison unit 310, configured to: After obtaining the fine-tuned retail store facade photo comparison model, when receiving an actual retail store facade photo taken by a user, acquire the current standard retail store facade photo of the retail store where the user is located, input the actual retail store facade photo and the current standard retail store facade photo into the fine-tuned retail store facade photo comparison model, and output a structured feature vector of the actual retail store facade photo and a structured vector of the current standard retail store facade photo, wherein the structured feature vector of the actual retail store facade photo includes the content of the facade area and the content of the fixed area around the facade of the actual retail store facade photo, and the structured vector of the current standard retail store facade photo includes the content of the facade area and the content of the fixed area around the facade of the current standard retail store facade photo. Calculate the similarity between the actual retail store facade photo and the current standard retail store facade photo according to the structured feature vector of the actual retail store facade photo and the structured vector of the current standard retail store facade photo; Determine whether the actual retail store facade photo and the current standard retail store facade photo are facade photos of the same retail store according to the similarity between the actual retail store facade photo and the current standard retail store facade photo.

[0099] In an embodiment of the present application, a possible implementation manner is provided, and the comparison unit 310 is further configured to: If the similarity between the actual retail store facade photo and the current standard retail store facade photo is greater than a preset threshold, determine that the actual retail store facade photo and the current standard retail store facade photo are facade photos of the same retail store; If the similarity between the actual retail store facade photo and the current standard retail store facade photo is less than or equal to the preset threshold, determine that the actual retail store facade photo and the current standard retail store facade photo are not facade photos of the same retail store.

[0100] In an embodiment of the present application, a possible implementation manner is provided, and the second acquisition unit 220 is further configured to: Collect retail store facade photos in multiple situations, including one or more of the following: Collect photos of the same retail store from different angles and at different distances; Collect photos of the same retail store with vehicles in front of the store; Collect photos of the same retail store with pedestrians in front of the store; Collect photos of the same retail store with no obstruction in front of the store; Collect photos of the same retail store where the proportion of the store facade in the photo is within a set proportion range.

[0101] In an embodiment of the present application, a possible implementation manner is provided. When collecting photos of the same retail store from different angles and at different distances, the data obtained by performing structured annotation on the retail store facade photos in multiple situations for the facade area position and content, the fixed area position and content around the facade, and the obstacle area position and content further includes descriptions of the facade angle and distance; When collecting photos of the same retail store with vehicles in front of the store, the data obtained by performing structured annotation on the retail store facade photos in multiple situations for the facade area position and content, the fixed area position and content around the facade, and the obstacle area position and content further includes a description of the position of the vehicle relative to the facade; When collecting photos of pedestrians in front of the same retail store, the data obtained by structuring and annotating the photos of the storefront of the retail store in the above-mentioned multiple situations for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of the position of the pedestrian relative to the storefront. When collecting unobstructed photos in front of the same retail store, the data obtained by structuring and annotating the photos of the storefront of the retail store in the above-mentioned multiple situations for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of no obstruction in front of the store. When collecting photos in which the proportion of the storefront of the same retail store in the photo is within a set proportion range, the data obtained by structuring and annotating the photos of the storefront of the retail store in the above-mentioned multiple situations for the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area also includes a description of the proportion of the storefront in the photo.

[0102] In an embodiment of the present application, a possible implementation manner is provided, and the fine-tuning training unit 230 is further configured to: Load the Low-Rank Adaptation (LoRA) module to perform fine-tuning training on the preset multi-modal model. During the fine-tuning training, insert the first weight corresponding to the position and content of the storefront area, the second weight corresponding to the position and content of the fixed area around the storefront, and the third weight corresponding to the position and content of the obstacle area, update the parameters of the LoRA module, and freeze the main parameters of the preset multi-modal model to obtain a fine-tuned retail storefront photo comparison model.

[0103] In an embodiment of the present application, a possible implementation manner is provided. In the data obtained by structuring and annotating the position and content of the storefront area, the position and content of the fixed area around the storefront of the standard retail storefront photo, the position of the storefront area and the position of the fixed area around the storefront are respectively represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the standard retail storefront photo. In the data obtained by structuring and annotating the position and content of the storefront area, the position and content of the fixed area around the storefront, and the position and content of the obstacle area of the retail storefront photos in the above-mentioned multiple situations, the position of the storefront area, the position of the fixed area around the storefront, and the position of the obstacle area are respectively represented by rectangular coordinates, and the coordinate values are the number of pixels starting from the upper left corner of the retail storefront photos in the above-mentioned multiple situations.

[0104] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, including a processor and a memory. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the training method of the retail storefront photo comparison model based on the multi-modal model in any one of the above embodiments.

[0105] In an exemplary embodiment, an electronic device is provided, such as Figure 4 shown Figure 4 the electronic device 400 shown in FIG. 400 includes: a processor 401 and a memory 403. Among them, the processor 401 and the memory 403 are connected, such as connected through a bus 402. Optionally, the electronic device 400 may further include a transceiver 404. It should be noted that in practical applications, the transceiver 404 is not limited to one, and the structure of the electronic device 400 does not constitute a limitation to the embodiments of the present application.

[0106] The processor 401 may be a CPU (Central Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), FPGA or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present application. The processor 401 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0107] The bus 402 may include a path for transmitting information between the above components. The bus 402 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 402 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is used to represent it in FIG., but it does not mean that there is only one bus or one type of bus.

[0108] The memory 403 may be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0109] The memory 403 is used to store the computer program code for executing the solution of this application and is controlled by the processor 401 for execution. The processor 401 is used to execute the computer program code stored in the memory 403 to implement the content shown in the foregoing method embodiments.

[0110] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0111] Based on the same inventive concept, the embodiments of this application also provide a storage medium in which a computer program is stored. The computer program is set to execute the training method of the retail store facade photo comparison model based on the multi-modal model in any one of the foregoing embodiments when running.

[0112] Those skilled in the art can clearly understand the specific working processes of the above-described systems, devices, and modules. They can refer to the corresponding processes in the foregoing method embodiments. For the sake of brevity, they will not be described in detail here.

[0113] Those of ordinary skill in the art can understand that the technical solution of this application can essentially or in whole or in part be embodied in the form of a software product. This computer software product is stored in a storage medium, which includes a number of program instructions for causing an electronic device (such as a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application when the program instructions are running. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc.

[0114] Alternatively, all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions (such as an electronic device such as a personal computer, a server, or a network device, etc.). The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the electronic device, the electronic device executes all or part of the steps of the methods described in the various embodiments of this application.

[0115] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that within the spirit and principles of this application, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the protection scope of this application.

Claims

1. A training method for a retail store front photo comparison model based on a multimodal model, characterized in that: The method enables the retail store front photo comparison model after fine-tuning training to focus on the location and content of the front area and the location and content of the fixed area around the front. The method includes: Obtain a standard retail store front photo, and obtain structured annotation data of the store front area location and content, and the location and content of a fixed area around the store front for the standard retail store front photo; Collect photos of retail store fronts in various situations, and obtain data for structured annotation of the store front area location and content, the location and content of the fixed area around the store front, and the location and content of the obstacle area for the photos of retail store fronts in various situations; The standard retail store front photos, the data of structured annotation of the store front area position and content, and the position and content of the fixed area around the store front for the standard retail store front photos, the retail store front photos in the multiple situations, and the data of structured annotation of the store front area position and content, the position and content of the fixed area around the store front, and the position and content of the obstacle area for the retail store front photos in the multiple situations are input into a preset multimodal model, and the preset multimodal model is fine-tuned and trained. During the fine-tuning training, a first weight corresponding to the store front area position and content, a second weight corresponding to the fixed area position and content around the store front, and a third weight corresponding to the obstacle area position and content are inserted to obtain a retail store front photo comparison model after fine-tuning and training; wherein the third weight is less than the first weight, and the third weight is less than the second weight.

2. The method according to claim 1, characterized in that After obtaining the retail store front photo comparison model after fine-tuning training, the method further includes: When receiving an actual retail store front photo taken by a user, obtain a current standard retail store front photo of the retail store where the user is located, input the actual retail store front photo and the current standard retail store front photo into the retail store front photo comparison model after fine-tuning training, and output a structured feature vector of the actual retail store front photo and a structured vector of the current standard retail store front photo, wherein the structured feature vector of the actual retail store front photo contains the content of the door front area and the content of the fixed area around the door front of the actual retail store front photo, and the structured vector of the current standard retail store front photo contains the content of the door front area and the content of the fixed area around the door front of the current standard retail store front photo; Calculating the similarity between the actual retail store front photo and the current standard retail store front photo based on the structured feature vector of the actual retail store front photo and the structured vector of the current standard retail store front photo; According to the similarity between the actual retail store door front photo and the current standard retail store door front photo, it is determined whether the actual retail store door front photo and the current standard retail store door front photo are door front photos of the same retail store.

3. The method according to claim 2, characterized in that According to the similarity between the actual retail store front photo and the current standard retail store front photo, determining whether the actual retail store front photo and the current standard retail store front photo are the same retail store front photo, including: If the similarity between the actual retail store front photo and the current standard retail store front photo is greater than a preset threshold, it is determined that the actual retail store front photo and the current standard retail store front photo are photos of the same retail store; If the similarity between the actual retail store front photo and the current standard retail store front photo is less than or equal to a preset threshold, it is determined that the actual retail store front photo and the current standard retail store front photo are not photos of the same retail store.

4. The method according to claim 1, characterized in that: Collect photos of retail storefronts in a variety of situations, including one or more of the following: Collect photos of the same retail store from different angles and distances; Collect photos of vehicles in front of the same retail store; Collect photos of people walking in front of the same retail store; Collect unobstructed photos of the same retail store front; Collect photos of the same retail store with the storefront occupying a certain proportion within the set range.

5. The method according to claim 4, characterized in that When collecting photos of the same retail store at different angles and distances, the data for obtaining structured annotations of the storefront area location and content, the fixed area location and content around the storefront, and the obstacle area location and content for the photos of the storefront in the various situations also includes descriptions of the storefront angle and distance; When collecting photos of vehicles in front of the same retail store, the data obtained for structured annotation of the storefront area position and content, the fixed area position and content around the storefront, and the obstacle area position and content for the photos of the retail storefront in the above multiple situations also includes a description of the position of the vehicle relative to the storefront; When collecting photos of pedestrians in front of the same retail store, the data obtained for structured annotation of the storefront area position and content, the fixed area position and content around the storefront, and the obstacle area position and content for the photos of the retail storefront in the above multiple situations also includes a description of the position of the pedestrian relative to the storefront; When collecting photos of the same retail store front without obstructions, the data obtained for structured annotation of the retail store front photos in the above multiple situations for the location and content of the front area, the location and content of the fixed area around the front, and the location and content of the obstacle area also includes a description of the unobstructed front of the door; When collecting photos of the same retail store front in which the ratio of the store front to the photo is within a set ratio range, the data obtained for structured annotation of the store front area location and content, the fixed area location and content around the store front, and the obstacle area location and content for the photos of the retail store front in the multiple cases also includes a description of the ratio of the store front to the photo.

6. The method according to any one of claims 1 to 5, characterized in that Fine-tune the preset multimodal model, and insert a first weight corresponding to the location and content of the door head area, a second weight corresponding to the location and content of the fixed area around the door head, and a third weight corresponding to the location and content of the obstacle area during the fine-tune training, to obtain a retail store door head photo comparison model after fine-tuning training, including: Load the low-rank adaptive LoRA module to perform fine-tuning training on the preset multimodal model. During the fine-tuning training, insert the first weight corresponding to the position and content of the doorhead area, the second weight corresponding to the position and content of the fixed area around the doorhead, and the third weight corresponding to the position and content of the obstacle area, update the parameters of the LoRA module, freeze the main parameters of the preset multimodal model, and obtain the retail store doorhead photo comparison model after fine-tuning training.

7. The method according to any one of claims 1 to 5, characterized in that In the data of structured annotation of the storefront area position and content, and the fixed area position and content around the storefront for the standard storefront photo, the storefront area position and the fixed area position around the storefront are respectively represented by rectangular coordinates, and the coordinate value is the number of pixels starting from the upper left corner of the standard storefront photo; In the data of structured annotation of the storefront area position and content, the fixed area position and content around the storefront, and the obstacle area position and content of the retail storefront photos in the multiple situations described above, the storefront area position, the fixed area position around the storefront, and the obstacle area position are respectively represented by rectangular coordinates, and the coordinate value is the number of pixels starting from the upper left corner of the retail storefront photos in the multiple situations described above.

8. A training device for a retail store front photo comparison model based on a multimodal model, characterized in that: The device enables the retail store front photo comparison model after fine-tuning training to focus on the position and content of the front area and the position and content of the fixed area around the front, and the device includes: A first acquisition unit is used to acquire a standard retail store front photo, and acquire data for structured annotation of the store front area position and content, and the fixed area position and content around the store front for the standard retail store front photo; A second acquisition unit is used to collect photos of retail store fronts in various situations, and obtain data for structured annotation of the store front area position and content, the fixed area position and content around the store front, and the obstacle area position and content for the photos of retail store fronts in various situations; A fine-tuning training unit is used to input the standard retail store front photos, the data of structured annotation of the store front area position and content, and the position and content of the fixed area around the store front for the standard retail store front photos, the retail store front photos in the multiple situations, and the data of structured annotation of the store front area position and content, the position and content of the fixed area around the store front, and the position and content of the obstacle area for the retail store front photos in the multiple situations, into a preset multimodal model, fine-tune the preset multimodal model, and insert a first weight corresponding to the store front area position and content, a second weight corresponding to the fixed area position and content around the store front, and a third weight corresponding to the obstacle area position and content during the fine-tuning training to obtain a retail store front photo comparison model after fine-tuning training; wherein the third weight is less than the first weight, and the third weight is less than the second weight.

9. An electronic device, characterized in that: The invention comprises a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the training method of the retail store front photo comparison model based on the multimodal model as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method for training a retail store front photo comparison model based on a multimodal model according to any one of claims 1 to 7 when running.

Citation Information

Patent Citations

  • Marking method of background of marked image and skin problem identification method and device

    CN114022936A

  • Cigarette delivery destination consistency inspection system, method, equipment and medium

    CN117809066A

  • Electronic device, method for constructing scoring model of retail outlets, system, and computer readable medium

    US20210125131A1