Urban governance multi-modal large model construction method based on target area

By building a multimodal model of urban governance, using visual and text feature extractors and language models to generate detailed scene descriptions, the problem that existing models cannot accurately identify illegal scenarios is solved, and more accurate urban management is achieved.

CN120472292APending Publication Date: 2025-08-12CHENGDU ZHIHUI HENENG CITY TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510365952.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing urban management model cannot accurately distinguish violations in urban management, such as outdoor hanging clothes and outside store business problems, and cannot give accurate descriptions and positioning.

Method used

Build a multimodal model of urban governance based on the target area. By obtaining the training set, establish and train a multimodal model of urban governance. Use visual and text feature extractors, language models and multimodal model to generate detailed scene descriptions and qualitative judgments and select the optimal model.

Benefits of technology

It realizes more accurately identifying violation scenarios in urban management, provides detailed scenario descriptions and qualitative judgments, and improves the accuracy of urban governance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472292A_ABST
    Figure CN120472292A_ABST
Patent Text Reader

Abstract

The invention discloses an urban governance multi-modal large model construction method based on a target area, and relates to the technical field of smart cities. The method comprises the following steps: S1, acquiring a training set; s2, establishing an urban governance multi-modal large model; s3, training the urban governance multi-modal large model in the step S2 through the training set obtained in the step S1; a trained urban governance multi-modal large model is obtained; s4, repeatedly executing the step S3 to obtain a first number of trained urban governance multi-modal large models; and S5, evaluating the first number of trained urban governance multi-modal large models, and selecting the optimal urban governance multi-modal large model as the urban governance multi-modal large model constructed by the method. According to the urban governance multi-modal large model constructed by the urban governance multi-modal large model construction method based on the target area, violation scenes in urban management can be identified more accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart city technology, and more specifically, to a method for constructing a large multimodal model of urban governance based on a target area. Background Art

[0002] Current urban management models can only provide basic descriptions of perceived urban issues, but are unable to pinpoint specific urban issues or provide rich descriptions of them. For example, if laundry is hung on the street, this violation falls under the "outdoor hanging" (hanging clothes) violation under urban management; if laundry is hung outside a clothing store, this violation falls under the "out-of-store operation" (operational violation) violation under urban management. Existing models can only identify the presence of laundry hanging, but cannot distinguish between outdoor hanging and out-of-store operation violations. Summary of the Invention

[0003] The technical problem to be solved by this application is: how to more accurately identify violation scenarios in urban management.

[0004] The technical solution adopted by this application to solve its technical problems is:

[0005] In a first aspect, the present application provides a method for constructing a large multimodal model of urban governance based on a target area, the method comprising the following steps:

[0006] S1, obtain the training set;

[0007] S2, building a large multimodal model of urban governance;

[0008] S3, training the urban governance multimodal large model in step S2 using the training set obtained in step S1; obtaining a trained urban governance multimodal large model;

[0009] S4, repeating step S3 to obtain a first number of trained urban governance multimodal large models;

[0010] S5. Evaluate the first number of trained urban governance multimodal big models, and select the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction method based on the target area.

[0011] Based on the first aspect, further, the training set includes: a second picture, an instruction text, and a third text label.

[0012] Based on the first aspect, further, the third text label is obtained by using the following method:

[0013] S11, obtaining a first image, identifying the first image using an urban governance target detection model, and generating a first text label and a second image; the first text label includes a target category; and the second image includes a target positioning frame;

[0014] S12: Based on the first text label, the second image is understood by the universal multimodal large model to generate a second text label, where the second text label includes a second scene description, and the scene description is diverse.

[0015] S13, based on the second text label, the general multimodal large model is used to understand the area within the target positioning box in the second image, and generate a third text label, wherein the content of the third text label includes a third scene description, the third scene description includes a qualitative judgment result, and the third scene description has general logic.

[0016] Based on the first aspect, further, the specific method of generating the second text label by understanding the second image through the universal multimodal large model based on the first text label includes:

[0017] The first text label, the second picture and the first guide word are input into the universal multimodal large model, and the universal multimodal large model outputs the second text label.

[0018] Based on the first aspect, further, the specific method of generating the third text label based on the second text label by understanding the area within the target positioning box in the second image through the universal multimodal large model includes:

[0019] The second guide word is input into the universal multimodal large model; the universal multimodal large model combines the second image and the second text label generated in step S12 to output a third text label.

[0020] Based on the first aspect, further, the urban governance multimodal large model includes a visual space feature extractor, a text space converter, a first large language model segmenter, a first large language model embedding module and a first large language model encoding module; the specific method of training the urban governance multimodal large model in step S2 with the training set obtained in step S1 includes: inputting the second picture in the training set into the visual space feature extractor of the urban governance multimodal large model, inputting the output of the visual space feature extractor into the text space converter, and the text space converter outputs the picture embedding feature; inputting the instruction text in the training set into the first large language model segmenter, and the first large language model segmenter outputs the instruction text. Let the text be marked; input the instruction text mark into the first large language model embedding module, and the first large language model embedding module outputs the instruction text embedding feature; input the third text label in the training set into the first large language model word segmenter, and the first large language model word segmenter outputs the third text mark; input the third text mark into the first large language model embedding module, and the first large language model embedding module outputs the third text embedding feature; input the image embedding feature and the instruction text embedding feature into the first large language model encoding module, and the first large language model encoding module outputs the predicted text embedding feature; calculate the joint loss, and adjust the parameters of the urban governance multimodal large model so that the joint loss meets the first condition.

[0021] The first condition is that the joint loss is less than a first threshold.

[0022] Based on the first aspect, further, the calculation formula of the joint loss is:

[0023]

[0024] in, Indicates joint loss; Represents the KL divergence in vector space; represents the target recognition loss value in the text space; λ represents the target recognition loss value scaling factor;

[0025] Among them, the KL divergence in the vector space The calculation formula is:

[0026]

[0027] P is the probability distribution of the predicted text embedding feature of the urban governance multimodal large model; T is the probability distribution of the third text embedding feature; i is the index value of the embedding feature;

[0028] Among them, the target recognition loss value in the text space is The calculation formula is:

[0029]

[0030] Among them, lab p Predicting text labels for targets in multimodal large models of urban governance; lab t is the text label of the target in the third text label; Dis is lab p and lab t The text editing distance between them; IOU is the intersection-over-union ratio of the predicted target positioning box of the urban governance multimodal large model and the target positioning box in the second picture; j is the index value of the target in the text space.

[0031] Based on the first aspect, further, the method of evaluating the first number of trained urban governance multimodal big models and selecting the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction method based on the target area includes:

[0032] Calculate the comprehensive similarity score of each trained urban governance multimodal big model, and take the trained urban governance multimodal big model with the highest comprehensive similarity score as the optimal urban governance multimodal big model.

[0033] Based on the first aspect, further, the calculation formula of the above-mentioned similarity comprehensive score is:

[0034]

[0035] Among them, Score represents the comprehensive score of similarity; K represents lab p and lab t Jaccard coefficient; Text p Represents the predicted text of the multimodal large model of urban governance; Text t Indicates the third text label; Sim llm Represents the semantic similarity score given by the second largest language model.

[0036] In the second aspect, the present application provides a system for constructing a multimodal large model of urban governance based on a target area, the system comprising a training set acquisition module, a model building module, a model training module, and a model evaluation module, wherein:

[0037] A training set acquisition module is used to obtain a training set;

[0038] Model building module, used to build a large multimodal model of urban governance;

[0039] A model training module is used to train the urban governance multimodal large model established by the model establishment module using the training set obtained by the training set acquisition module; and obtain a first number of trained urban governance multimodal large models;

[0040] The model evaluation module is used to evaluate the first number of trained urban governance multimodal big models and select the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction system based on the target area.

[0041] In a third aspect, the present application also provides an electronic device comprising a memory and a processor; the memory is used to store one or more programs; when the one or more programs are executed by the processor, a method as described in any one of the first aspects above is implemented.

[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements any one of the methods in the first aspect above when executed by a processor.

[0043] The beneficial effects of this application are:

[0044] The multimodal large model of urban governance constructed by the method for constructing a multimodal large model of urban governance based on the target area in this application can more accurately identify violation scenarios in urban management. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A flowchart of the method for constructing a multimodal large model of urban governance provided in an embodiment of the present application.

[0046] Figure 2 This is a principle block diagram of the urban governance multimodal large model construction system provided in the embodiment of the present application.

[0047] Figure 3 This is an example of the first picture provided in the embodiment of the present application.

[0048] Figure 4 This is another example of the first picture provided in the embodiment of the present application.

[0049] Figure 5 This is an example of the second picture provided in the embodiment of the present application.

[0050] Figure 6 This is another example of the second picture provided in the embodiment of the present application.

[0051] Figure 7 This is a structural block diagram of the electronic device provided in an embodiment of the present application.

[0052] The accompanying drawings are marked as follows: 100, training set acquisition module; 200, model building module; 300, model training module; 400, model evaluation module; 101, memory; 102, processor; 103, communication interface. DETAILED DESCRIPTION

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0054] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0055] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0056] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0057] In the description of the embodiments of the present application, "a plurality of" means at least 2.

[0058] First, as Figure 1 As shown, an embodiment of the present application provides a method for constructing a multimodal large model of urban governance based on a target area, the method comprising the following steps:

[0059] S1, obtain the training set;

[0060] S2, building a large multimodal model of urban governance;

[0061] S3, training the urban governance multimodal large model in step S2 using the training set obtained in step S1; obtaining a trained urban governance multimodal large model;

[0062] S4, repeating step S3 to obtain a first number of trained urban governance multimodal large models;

[0063] S5. Evaluate the first number of trained urban governance multimodal big models, and select the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction method based on the target area.

[0064] In one embodiment of the present application, the training set includes: a second picture, an instruction text, and a third text label.

[0065] In one embodiment of the present application, a specific method for obtaining the third text label includes the following steps:

[0066] S11, obtaining a first image, identifying the first image using an urban governance target detection model, and generating a first text label and a second image; the first text label includes a target category; and the second image includes a target positioning frame;

[0067] S12: Based on the first text label, the second image is understood by the universal multimodal large model to generate a second text label, where the second text label includes a second scene description, and the scene description is diverse.

[0068] S13, based on the second text label, the general multimodal large model is used to understand the area within the target positioning box in the second image, and generate a third text label, wherein the content of the third text label includes a third scene description, the third scene description includes a qualitative judgment result, and the third scene description has general logic.

[0069] The general logic mentioned above refers to general knowledge and logical judgment ability.

[0070] An example of the first picture is as follows Figure 3 shown.

[0071] Another example of the first picture is Figure 4 shown.

[0072] In one embodiment of the present application, the urban governance target detection model refers to a small and medium-sized target detection model represented by the Yolo series.

[0073] In one embodiment of the present application, the first picture is a picture of a scene of urban governance violations.

[0074] In one embodiment of the present application, the second picture is a picture with a target positioning frame.

[0075] In one embodiment of the present application, the universal multimodal large model may be GLM-4V or Qwen2.5-VL.

[0076] In one embodiment of the present application, the content of the third text tag includes the content of the second text tag, and the content of the second text tag includes the content of the first text tag.

[0077] In one embodiment of the present application, the specific method of generating the second text label by understanding the second image through the universal multimodal large model based on the first text label includes:

[0078] The first text label, the second picture and the first guide word are input into the universal multimodal large model, and the universal multimodal large model outputs the second text label.

[0079] In one embodiment of the present application, the specific method of generating the third text label by understanding the area within the target positioning box in the second image using the universal multimodal large model based on the second text label includes:

[0080] The second guide word is input into the universal multimodal large model; the universal multimodal large model combines the second image and the second text label generated in step S12 to output a third text label.

[0081] like Figure 5 As shown, in one embodiment of the present application, the red box in the second image is the target positioning box. The first text label reads: "There is clothing placed outdoors in the image [[631,306,728,445]]." 631 is the horizontal coordinate of the upper left corner of the target positioning box, 306 is the vertical coordinate of the upper left corner of the target positioning box, 728 is the horizontal coordinate of the lower right corner of the target positioning box, and 445 is the vertical coordinate of the lower right corner of the target positioning box. "Clothing placed outdoors" is the target category.

[0082] The first guide word could be: "Describe the content of the picture, focusing on the area [[631,306,728,445]] in the picture."

[0083] The first guide word can also be: "Describe the content of the picture, and focus on the red box area in the picture."

[0084] The content of the second text tag is: "This is a photo of a city street scene. A street with buildings and shops on both sides can be seen in the photo. Several people are walking on the street, and there are some trees on both sides of the street. The outdoor clothing display area [[631,306,728,445]] is a clothing store. Several clothes are displayed outside the store, which may be a promotional activity to attract customers' attention." Among them, "This is a photo of a city street scene. A street with buildings and shops on both sides can be seen in the photo. Several people are walking on the street, and there are some trees on both sides of the street. The outdoor clothing display area [[631,306,728,445]] is a clothing store. Several clothes are displayed outside the store, which may be a promotional activity to attract customers' attention" describes the second scene.

[0085] The second guide is: "Based on the output just now and based on the area where the red box is located, explain the urban management issues and assign a severity level ('severe', 'normal', 'minor') based on their impact."

[0086] The content of the third text label is: "This is a photo of a city street scene. A street can be seen in the photo, with buildings and shops on both sides. There are several people walking on the street, and there are some trees on both sides of the street. The outdoor clothing display area [[631,306,728,445]] is a clothing store. Several pieces of clothing are displayed outside the store. It may be a promotional activity to attract customers' attention. According to the area [[631,306,728,445]], it shows that a clothing store displays and sells clothes outside the door. In terms of urban management, this is an off-store operation problem. Placing goods on the sidewalk may occupy public space and affect pedestrian traffic, especially during peak hours, which may cause traffic congestion. This problem can be classified as a 'general' level of severity. Although it does have certain negative effects, it can usually be solved by strengthening management and law enforcement, and will not cause serious social problems. Disruption of public order or other major consequences. However, long-term neglect of such issues may gradually accumulate into larger management challenges, so they still require attention and appropriate measures to regulate them. The third scenario describes, "According to areas [[631,306,728,445]], a clothing store displays and sells clothing outside its doors. From an urban management perspective, this constitutes an off-site operation. Placing merchandise on the sidewalk may occupy public space, hinder pedestrian traffic, and may cause traffic congestion, especially during peak hours. This issue can be classified as 'moderate' in severity. While it does have certain negative impacts, it can usually be resolved through strengthened management and law enforcement, without causing serious disruption of public order or other major consequences. However, long-term neglect of such issues may gradually accumulate into larger management challenges, so they still require attention and appropriate measures to regulate them." "Outdoor display of clothing" is the text label of the target in the third text label.

[0087] The instruction text is: Describe the content of the picture and point out the problems in urban management.

[0088] like Figure 6 As shown in another embodiment of the present application, the red box in the second image is a target location box. The first text label reads: "There is clothing placed outdoors in the image [[341, 229, 578, 466]]." Here, 341 is the horizontal coordinate of the upper left corner of the target location box, 229 is the vertical coordinate of the upper left corner of the target location box, 578 is the horizontal coordinate of the lower right corner of the target location box, and 466 is the vertical coordinate of the lower right corner of the target location box. "Clothing placed outdoors" is the target category.

[0089] The first guide word could be: "Describe the content of the picture, focusing on the area [[341,229,578,466]] in the picture."

[0090] The first guide word can also be: "Describe the content of the picture, and focus on the red box area in the picture."

[0091] The content of the second text label is: "This picture shows a narrow street with old brick buildings on both sides. There are several cars parked on the street and some trash cans on the roadside. The outdoor clothing area [[341,229,578,466]] is some clothes hanging on the rope. These clothes include clothes of different colors, such as red, blue, etc. They are hung on a rope spanning two houses and dried in the sun." Among them, "This picture shows a narrow street with old brick buildings on both sides. There are several cars parked on the street and some trash cans on the roadside. The outdoor clothing area [[341,229,578,466]] is some clothes hanging on the rope. These clothes include clothes of different colors, such as red, blue, etc. They are hung on a rope spanning two houses and dried in the sun" is a description of the second scene.

[0092] The second guide is: "Based on the output just now and based on the area where the red box is located, explain the urban management issues and assign a severity level ('severe', 'normal', 'minor') based on their impact."

[0093] The content of the third text tag is: "This picture shows a narrow street with old brick buildings on both sides. There are several cars parked on the street and some trash cans on the roadside. The outdoor clothing area [[341,229,578,466]] is some clothes hanging on the rope. These clothes include clothes of different colors, such as red, blue, etc. They are hung on a rope spanning two houses and dried in the sun. According to the situation in the picture, we can see that some clothes are hung directly on the rope in public places. In terms of urban management, this is an outdoor hanging problem. This behavior It may have a certain negative impact on urban management and environmental sanitation. From an aesthetic point of view, hanging clothes randomly may affect the overall image and cleanliness of the city. If such phenomena are widespread and unmanaged, it may lead to more people imitating this practice, thus forming a bad social atmosphere. I think this problem belongs to the "general" level. Although it will not cause serious consequences immediately, if it is not resolved or ignored for a long time, then over time, its harmfulness will gradually increase. Therefore, it is recommended that relevant departments should strengthen the supervision of such behaviors and take effective measures. Effective measures can be taken to guide citizens to develop good living habits and a sense of social morality. Public awareness of environmental protection can also be raised through publicity and education, allowing us to jointly protect the environment we live in. The third scenario describes, "Based on the situation in the picture, we can see some clothes hanging directly on ropes in public places. From an urban management perspective, this constitutes outdoor hanging. This behavior may have a certain negative impact on urban management and environmental sanitation. From an aesthetic perspective, randomly hanging clothes can affect the overall image and cleanliness of a city. If this phenomenon is widespread and unmanaged, it may lead to more people following suit, fostering a negative social atmosphere. I consider this issue to be 'general'. While it may not have immediate serious consequences, if it is left unaddressed or ignored for a long time, its harmful effects will gradually increase over time. Therefore, it is recommended that relevant departments strengthen supervision of such behavior and take effective measures to guide citizens to develop good living habits and a sense of social morality. Public awareness of environmental protection can also be raised through publicity and education, allowing us to jointly protect the environment we live in." "Outdoor hanging of clothes" is the text label for the target in the third text label.

[0094] The instruction text is: Describe the content of the picture and point out the problems in urban management.

[0095] In one embodiment of the present application, the urban governance multimodal large model includes a visual space feature extractor, a text space converter, a first large language model segmenter, a first large language model embedding module and a first large language model encoding module; the specific method of training the urban governance multimodal large model in step S2 with the training set obtained in step S1 includes: inputting the second picture in the training set into the visual space feature extractor of the urban governance multimodal large model, inputting the output of the visual space feature extractor into the text space converter, and the text space converter outputs the picture embedding feature; inputting the instruction text in the training set into the first large language model segmenter, and the first large language model segmenter outputs the instruction text. Let the text be marked; input the instruction text mark into the first large language model embedding module, and the first large language model embedding module outputs the instruction text embedding feature; input the third text label in the training set into the first large language model word segmenter, and the first large language model word segmenter outputs the third text mark; input the third text mark into the first large language model embedding module, and the first large language model embedding module outputs the third text embedding feature; input the image embedding feature and the instruction text embedding feature into the first large language model encoding module, and the first large language model encoding module outputs the predicted text embedding feature; calculate the joint loss, and adjust the parameters of the urban governance multimodal large model so that the joint loss meets the first condition.

[0096] In one embodiment of the present application, the visual spatial feature extractor is a Vision Transformer, which can be either ViT-bigG or EVA2-CLIP-E. The Vision Transformer is a model that introduces the Transformer architecture into the field of computer vision. The visual spatial feature extractor can divide an image into multiple small blocks and then convert these small blocks into vector sequences. The visual spatial feature extractor then processes these vector sequences using a Transformer (a deep learning model) encoder to extract high-level features of the image.

[0097] In one embodiment of the present application, the text space converter is a two-layer MLP (i.e., multi-layer perceptron) network. The text space converter can convert data from different sources or different modalities into a text feature space.

[0098] The first large language model includes a first large language model word segmenter, a first large language model embedding module and a first large language model encoding module. In one embodiment of the present application, the first large language model can be Qwen2.5 or GLM-4.

[0099] Image embedding features refer to the feature vectors used to represent images.

[0100] The instruction text markup refers to a series of "tokens" corresponding to the instruction text. These tokens can be words, subwords, characters, or a combination thereof.

[0101] In one embodiment of the present application, the instruction text tag is obtained by inputting the instruction text into a first large language model word segmenter, and the first large language model word segmenter decomposing the instruction text.

[0102] The third text mark refers to a series of “tokens” corresponding to the third text label.

[0103] In one embodiment of the present application, the third text tag is obtained by inputting the third text tag into the first large language model word segmenter, and the first large language model word segmenter decomposing the third text tag.

[0104] Predicted text embedding features refer to the feature vectors used to represent predicted text.

[0105] The first condition may be that the joint loss is less than a first threshold.

[0106] The first threshold may be preset.

[0107] The first threshold may be 0.5.

[0108] The calculation formula of the joint loss is:

[0109]

[0110] in, Indicates joint loss; Represents the KL divergence in vector space; represents the target recognition loss value in the text space; λ represents the target recognition loss value scaling factor, and its value range is [0,1];

[0111] Among them, the KL divergence in the vector space The calculation formula is:

[0112]

[0113] P is the probability distribution of the predicted text embedding feature of the urban governance multimodal large model; T is the probability distribution of the third text embedding feature; i is the index value of the embedding feature;

[0114] Among them, the target recognition loss value in the text space is The calculation formula is:

[0115]

[0116] Among them, labp Predicting text labels for targets in multimodal large models of urban governance; lab t is the text label of the target in the third text label; Dis is lab p and lab t The text editing distance between them; IOU is the intersection-over-union ratio of the predicted target positioning box of the urban governance multimodal large model and the target positioning box in the second picture; j is the index value of the target in the text space.

[0117] The predicted text can be obtained by embedding the predicted text into the first language model word segmenter and decoding it by the first language model word segmenter.

[0118] In one embodiment of the present application, the first number is 10.

[0119] In one embodiment of the present application, the specific method of evaluating a first number of trained urban governance multimodal big models and selecting the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction method based on the target area includes:

[0120] Calculate the comprehensive similarity score of each trained urban governance multimodal big model, and take the trained urban governance multimodal big model with the highest comprehensive similarity score as the optimal urban governance multimodal big model.

[0121] The calculation formula of the similarity comprehensive score is:

[0122]

[0123] Among them, Score represents the comprehensive score of similarity; K represents lab p and lab t Jaccard coefficient; Text p Represents the predicted text of the multimodal large model of urban governance; Text t Indicates the third text label; Sim llm It represents the semantic similarity score given by the second largest language model. The value range is [0,1]. The larger the value, the higher the semantic similarity.

[0124] In one embodiment of the present application, the second largest language model may be Qwen or GLM.

[0125] Second, as Figure 2 As shown, the embodiment of the present application provides a system for constructing a multimodal large model of urban governance based on a target area, the system including a training set acquisition module, a model building module, a model training module and a model evaluation module, wherein:

[0126] A training set acquisition module is used to obtain a training set;

[0127] Model building module, used to build a large multimodal model of urban governance;

[0128] A model training module is used to train the urban governance multimodal large model established by the model establishment module using the training set obtained by the training set acquisition module; and obtain a first number of trained urban governance multimodal large models;

[0129] The model evaluation module is used to evaluate the first number of trained urban governance multimodal big models and select the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction system based on the target area.

[0130] Thirdly, as Figure 7 As shown, an embodiment of the present application also provides an electronic device, which includes a memory 101 and a processor 102; the memory 101 is used to store one or more programs; when one or more programs are executed by the processor 102, a method as described in any one of the first aspects above is implemented.

[0131] The electronic device may further include a communication interface 103. The memory 101, processor 102, and communication interface 103 are electrically connected to each other directly or indirectly to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines. The memory 101 may be used to store software programs and modules, and the processor 102 executes the software programs and modules stored in the memory 101 to perform various functional applications and data processing. The communication interface 103 may be used to communicate signaling or data with other node devices.

[0132] Among them, the memory 101 can be but is not limited to: Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc.

[0133] The processor 102 may be an integrated circuit chip with signal processing capabilities. The processor 102 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0134] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by the processor 102, implements a method as described in any one of the first aspects above. If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0135] The above are merely preferred embodiments of the present application and are not intended to limit the present application. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

[0136] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A method for constructing a large multimodal model of urban governance based on a target area, characterized by: The following steps are involved: S1, obtain the training set; S2, building a large multimodal model of urban governance; S3, training the urban governance multimodal large model in step S2 using the training set obtained in step S1; obtaining a trained urban governance multimodal large model; S4, repeating step S3 to obtain a first number of trained urban governance multimodal large models; S5. Evaluate the first number of trained urban governance multimodal big models, and select the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction method based on the target area.

2. The method for constructing a large multimodal model of urban governance based on a target area according to claim 1 is characterized in that: The training set includes: a second picture, an instruction text, and a third text label.

3. The method for constructing a large multimodal model of urban governance based on a target area according to claim 2 is characterized in that: The third text label is obtained by the following method: S11, obtaining a first image, identifying the first image using an urban governance target detection model, and generating a first text label and a second image; the first text label includes a target category; and the second image includes a target positioning frame; S12: Based on the first text label, the second image is understood by the universal multimodal large model to generate a second text label, where the second text label includes a second scene description, and the scene description is diverse. S13, based on the second text label, the general multimodal large model is used to understand the area within the target positioning box in the second image, and generate a third text label, wherein the content of the third text label includes a third scene description, the third scene description includes a qualitative judgment result, and the third scene description has general logic.

4. The method for constructing a large multimodal model of urban governance based on a target area according to claim 3 is characterized in that: The specific method of generating the second text label by understanding the second image through the universal multimodal large model based on the first text label includes: The first text label, the second picture and the first guide word are input into the universal multimodal large model, and the universal multimodal large model outputs the second text label.

5. The method for constructing a large multimodal model of urban governance based on a target area according to claim 3 is characterized in that: The specific method of generating the third text label by understanding the area within the target positioning frame in the second image based on the second text label using the universal multimodal large model includes: The second guide word is input into the universal multimodal large model; the universal multimodal large model combines the second image and the second text label generated in step S12 to output a third text label.

6. The method for constructing a large multimodal model of urban governance based on a target area according to claim 1 is characterized in that: The urban governance multimodal model includes a visual space feature extractor, a text space converter, a first language model word segmenter, a first language model embedding module and a first language model encoding module; The specific method for training the urban governance multimodal large model in step S2 with the training set obtained in step S1 includes: inputting the second picture in the training set into the visual space feature extractor of the urban governance multimodal large model, inputting the output of the visual space feature extractor into the text space converter, and the text space converter outputs the picture embedding feature; inputting the instruction text in the training set into the first large language model word segmenter, and the first large language model word segmenter outputs the instruction text tag; inputting the instruction text tag into the first large language model embedding module, and the first large language model embedding module outputs the instruction text embedding feature; inputting the third text tag in the training set into the first large language model word segmenter, and the first large language model word segmenter outputs the third text tag; inputting the third text tag into the first large language model embedding module, and the first large language model embedding module outputs the third text embedding feature; inputting the picture embedding feature and the instruction text embedding feature into the first large language model encoding module, and the first large language model encoding module outputs the predicted text embedding feature; calculating the joint loss, and adjusting the parameters of the urban governance multimodal large model so that the joint loss meets the first condition.

7. The method for constructing a large multimodal model of urban governance based on a target area according to claim 6 is characterized in that: The first condition is that the joint loss is less than a first threshold.

8. The method for constructing a large multimodal model of urban governance based on a target area according to claim 6 is characterized in that: The calculation formula of the joint loss is: in, Indicates joint loss; Represents the KL divergence in vector space; represents the target recognition loss value in the text space; λ represents the target recognition loss value scaling factor; Among them, the KL divergence in the vector space The calculation formula is: P is the probability distribution of the predicted text embedding feature of the urban governance multimodal large model; T is the probability distribution of the third text embedding feature; i is the index value of the embedding feature; Among them, the target recognition loss value in the text space is The calculation formula is: Among them, lab p Predicting text labels for targets in multimodal large models of urban governance; lab t is the text label of the target in the third text label; Dis is lab p and lab t The text editing distance between them; IOU is the intersection-over-union ratio of the predicted target positioning box of the urban governance multimodal large model and the target positioning box in the second picture; j is the index value of the target in the text space.

9. The method for constructing a large multimodal model of urban governance based on a target area according to claim 1 is characterized in that: The specific method of evaluating the first number of trained urban governance multimodal big models and selecting the optimal urban governance multimodal big model as the urban governance multimodal big model constructed by the urban governance multimodal big model construction method based on the target area includes: Calculate the comprehensive similarity score of each trained urban governance multimodal big model, and take the trained urban governance multimodal big model with the highest comprehensive similarity score as the optimal urban governance multimodal big model.

10. The method for constructing a large multimodal model of urban governance based on a target area according to claim 9 is characterized in that: The calculation formula of the similarity comprehensive score is: Among them, Score represents the comprehensive score of similarity; K represents lab p and lab t Jaccard coefficient; Text p Represents the predicted text of the multimodal large model of urban governance; Text t Indicates the third text label; Sim llm Represents the semantic similarity score given by the second largest language model.

Citation Information

Cited By

  • Urban three-dimensional modeling and dynamic simulation environment generation method and device based on multi-modal large model

    CN120747380A

  • Method and device for generating a dynamic simulation environment based on a multimodal large model for a city three-dimensional modeling

    CN120747380B