Water-logged depth extraction method based on direct preference optimization of multi-modal large language model

By training a multimodal large language model using a token-level direct preference optimization strategy and combining text, image, and video modalities from social media data, the problem of accuracy in extracting water depth information during floods was solved, enabling real-time and detailed monitoring of water depth information.

CN119691495BActive Publication Date: 2026-02-03WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411559276.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2026-02-03
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately extract water depth information from complex social media data during floods, and traditional data sources cannot meet the needs of real-time monitoring.

Method used

A multimodal large language model trained based on a token-level direct preference optimization strategy is adopted, and social media data in text, image and video modalities are combined to perform water depth recognition through the multimodal large language model, and refined processing is performed through a water depth level model.

Benefits of technology

It enables accurate extraction of water depth information from social media data, improving the real-time performance and accuracy of flood disaster monitoring, and can obtain detailed water depth information based on target areas and time periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691495B_ABST
    Figure CN119691495B_ABST
Patent Text Reader

Abstract

The application provides a waterlogging water depth extraction method based on a direct preference optimized multi-modal large language model, comprising: obtaining social media content to be identified; inputting the social media content to be identified into a pre-trained water depth identification model to obtain a water depth identification result output by the water depth identification model, wherein the water depth identification model is a multi-modal large language model trained based on a token-level direct preference optimization strategy. The waterlogging water depth extraction method based on the direct preference optimized multi-modal large language model trains the multi-modal large language model through the token-level direct preference optimization strategy to obtain the water depth identification model, so that the water depth identification model can more accurately identify the water depth information corresponding to the input target region and target period of the social media content to be identified.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence application, and in particular to a waterlogging water depth extraction method based on direct preference optimization multi-modal large language model. BACKGROUND

[0002] Waterlogging or flood disaster caused by heavy rainfall is one of the most common disasters, and the waterlogging caused by flood disaster not only hinders traffic and causes congestion, but also causes damage to infrastructure such as streets and buildings. Collecting and identifying water depth information during the flood disaster process can directly reflect the severity of the flood disaster, thereby promoting the implementation of emergency measures, so it is very important to identify water depth during the flood disaster period.

[0003] Due to the limitation of meteorological factors and observation conditions, traditional waterlogging or flood disaster data sources are increasingly unable to meet the increasingly frequent flood disasters. When monitoring flood events in real time through airborne or remote sensing satellite images, poor weather conditions will limit the visibility of the images, especially during heavy rainfall. At the same time, the revisit time of the satellite also limits the timeliness of the data, and high-resolution optical remote sensing satellites usually need to obtain high-resolution images after the incident a few days later. Finally, through remote sensing images, it is impossible to obtain the waterlogging situation in underground areas such as garages and subways. Other methods also have certain problems, such as water level observation sensors cannot achieve large-scale monitoring due to limited distribution, and flood simulation models based on surface water and underground pipe networks require a large amount of computing resources and cannot achieve timely monitoring due to complex physical mechanisms.

[0004] Compared with traditional data sources, crowd-sourced social media big data can collect spatio-temporal information related to urban flood topics in real time. In the event of extreme disasters or emergencies, people will spontaneously seek help and upload disaster conditions and information through social media. By collecting social media big data, the development process of each stage of the disaster can be reflected, which is helpful for disaster prevention, emergency response and post-disaster recovery. Since the cost of collecting crowd-sourced social media big data is low, collecting people's reports from it to reflect the waterlogging area of the city is a feasible way to help improve the perception of urban flood depth information.

[0005] Therefore, how to accurately extract water depth information based on complex social media data is still a problem to be solved. SUMMARY

[0006] The present application provides a waterlogging water depth extraction method based on direct preference optimization multi-modal large language model, which solves the defect that it is difficult to accurately extract water depth information from complex social media data in the prior art, and realizes a waterlogging water depth extraction method based on social media data.

[0007] This invention provides a method for extracting water depth based on a direct preference-optimized multimodal large language model, comprising:

[0008] Obtain the social media content to be identified;

[0009] The social media content to be identified is input into a pre-trained depth recognition model to obtain the depth recognition result output by the model. The depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0010] According to the present invention, a method for extracting water depth based on a direct preference-optimized multimodal large language model, prior to the step of inputting the social media content to be identified into a pre-trained water depth recognition model, further includes:

[0011] The multimodal large language model was trained using a direct preference optimization strategy. During the training process, weights for different text modules were added to the original scoring calculation of the direct preference optimization strategy.

[0012] According to the present invention, a method for extracting water depth from a multimodal large language model based on direct preference optimization is provided, wherein before the step of training the multimodal large language model using the direct preference optimization strategy, the method further includes:

[0013] Obtain multiple social media data related to water depth and filter out social media data that includes geographical location;

[0014] Social media data of different modalities are labeled to obtain a dataset used to train the water depth recognition model.

[0015] According to the present invention, a method for extracting water depth based on a direct preference-optimized multimodal large language model is provided, wherein the step of annotating social media data of different modalities specifically includes:

[0016] For text modal data, extract and standardize the geographical location and water depth information from the text data according to a preset format;

[0017] For image modality data, multiple instructions for querying the image data are predetermined. The instructions and each piece of image data are sequentially input into the multimodal large language model to obtain the recognition result output by the multimodal large language model. Based on the recognition result, the error parts are corrected to obtain the correct response result. The correct response result is used as a positive sample, and the recognition result is used as a negative sample.

[0018] For video modal data, a frame of video data is extracted at preset intervals and labeled as image data.

[0019] According to the present invention, a method for extracting water depth based on a direct preference-optimized multimodal large language model is provided. After the step of obtaining the water depth recognition result output by the water depth recognition model, the method further includes:

[0020] Predefine water depth levels and classify water depth descriptions, and train a water depth level model based on the defined water depth levels;

[0021] The water depth identification result is input into the water depth classification model to obtain the water depth classification result output by the water depth classification model.

[0022] According to the present invention, a method for extracting water depth based on a direct preference-optimized multimodal large language model is provided. After the step of obtaining the water depth level results output by the water depth level model, the method further includes:

[0023] The water depth classification results are then visualized.

[0024] The present invention also provides a water depth extraction device based on a direct preference-optimized multimodal large language model, comprising:

[0025] The acquisition module is used to acquire the social media content to be identified;

[0026] The recognition module is used to input the social media content to be recognized into a pre-trained depth recognition model and obtain the depth recognition result output by the depth recognition model, wherein the depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0027] A water depth extraction device based on a direct preference-optimized multimodal large language model provided by the present invention further includes:

[0028] The grading module is used to predefine water depth grades and classify water depth descriptions, and to train a water depth grade model based on the defined water depth grades.

[0029] The water depth identification result is input into the water depth classification model to obtain the water depth classification result output by the water depth classification model.

[0030] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the water depth extraction method based on direct preference optimization of a multimodal large language model as described above.

[0031] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the water depth extraction method based on direct preference optimization of a multimodal large language model as described above.

[0032] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the water depth extraction method based on direct preference optimization of a multimodal large language model as described above.

[0033] The water depth extraction method based on direct preference optimization multimodal large language model provided by this invention trains the multimodal large language model with a token-level direct preference optimization strategy to obtain a water depth recognition model. This enables the water depth recognition model to more accurately identify the corresponding water depth information based on the social media content to be identified in the target area and target time period. Finally, it can obtain the water depth recognition result of the target area in the target time period based on multiple social media content to be identified in the target area and target time period. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0035] Figure 1 This is one of the flowcharts of the water depth extraction method based on direct preference optimization multimodal large language model provided by the present invention;

[0036] Figure 2 This is a schematic diagram mainly used to illustrate the optimization process of the token-level direct preference optimization strategy in the water depth extraction method based on direct preference optimization of multimodal large language models provided by this invention;

[0037] Figure 3 This is a schematic diagram used to illustrate the water depth recognition model in the water depth extraction method based on direct preference optimization multimodal large language model provided by the present invention;

[0038] Figure 4 This is a schematic diagram used to display water depth levels in the water depth extraction method based on direct preference optimization multimodal large language model provided by the present invention;

[0039] Figure 5 This is the second flowchart of the water depth extraction method based on direct preference optimization multimodal large language model provided by the present invention;

[0040] Figure 6 This is a schematic diagram of the water depth extraction device based on a direct preference-optimized multimodal large language model provided by the present invention.

[0041] Figure 7This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0043] The following is combined with Figures 1 to 4 This invention introduces a method for extracting water depth based on a direct preference-optimized multimodal large language model, such as... Figure 1 As shown, it includes:

[0044] Step 101: Obtain the social media content to be identified;

[0045] The social media content to be identified includes text, image, and video modalities, or combinations thereof.

[0046] Specifically, when it is necessary to identify the water depth of a target area at a target time, social media posts containing geographic location information and water depth information are obtained from different social media platforms as the social media content to be identified, thereby identifying the water depth information of the corresponding area based on the social media content to be identified.

[0047] Alternatively, social media can be a social media platform where text, images, or videos are the primary content posted.

[0048] Optionally, text content that directly contains geolocation information can be identified as social media content containing geolocation information, or social media content with location or tags can be identified as social media content containing geolocation information.

[0049] For example, if the text of the social media content to be identified is "There is standing water on Street A, it's up to my calves", it is considered to contain the geographical location information "Street A"; if the social media content to be identified contains the location information "Community B", it is considered to contain the geographical location information "Community B"; if the social media content to be identified contains the hashtag "#CDistrictNews", it is considered to contain the geographical location information "CDistrict C".

[0050] Optionally, if the text, images, and / or videos of social media content contain water depth information, then such content is identified as social media content containing water depth information to be identified.

[0051] Optionally, the image content in social media content includes both still images and moving images.

[0052] Optionally, the social media content to be identified can be one or more pieces.

[0053] Step 102: Input the social media content to be identified into the pre-trained depth recognition model to obtain the depth recognition result output by the depth recognition model, wherein the depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0054] Currently, most methods for extracting depth information based on social media data tend to develop single-modal models for a single data source (such as text or images). This leads to a waste of other modal data resources in social media data and also easily causes bias towards a single modality, ignoring important information carried in other modalities.

[0055] Therefore, in order to achieve water depth identification based on social media data more flexibly, so as to maximize the use of social media data to extract water depth information in a fine-grained manner, this embodiment uses a multimodal large language model as the basis for training to obtain a water depth identification model.

[0056] To fully utilize the diverse modalities of social media data, the selected multimodal large language model needs to be capable of handling large language model data across text, image, and video modalities. This implementation uses MiniCPM-V (an edge-side multimodal large language model) as the base model for illustration.

[0057] MiniCPM-V consists of a visual encoder, a compression layer, and a large language model. It is pre-trained on a Chinese corpus, enabling it to better adapt to social media content in the Chinese context. The visual encoder is responsible for encoding images and video frames, the compression layer is responsible for compressing the visually encoded tokens to a reasonable range, and the large language model is responsible for understanding and outputting the compressed visual and text encodings.

[0058] Although multimodal large language models provide a foundation for unified processing and encoding of data from different modalities, enabling unified understanding and processing by encoding and converting multimodal data into natural language, current large language models generally suffer from illusion problems, lacking sufficient understanding of water depth in images and easily outputting incorrect recognition results. Therefore, in order to unleash the model's ability to understand flood disasters and water levels and improve the model's accuracy, this implementation uses a token-level direct preference optimization strategy to train the selected multimodal large language model to obtain a water depth recognition model with more accurate recognition results.

[0059] The token-level direct preference optimization strategy differs from common direct preference optimization strategies. Instead, it refines the optimization process to the token level. Specifically, during the production response process, text is generated token by token, and the weights of the text segments corresponding to each token are adjusted to achieve more granular adjustment and optimization. This avoids overfitting of the model while fine-tuning the large language model.

[0060] Based on this, the social media content to be identified is input into the trained water depth recognition model, and the water depth recognition result output by the model can be obtained.

[0061] For example, if the text in the social media content to be identified is "The water has flooded my house" and contains the location "D community", and the image content is "The water has flooded the shoulder", then the water depth recognition model will generate a recognition result such as "The water depth in D community has submerged the shoulder".

[0062] Understandably, by collecting multiple pieces of social media content to be identified in the target area during the target time period and obtaining their water depth identification results, the water depth information of the target area during the target time period can be obtained based on the water depth identification results.

[0063] This invention trains a multimodal large language model to obtain a water depth recognition model through a token-level direct preference optimization strategy. This enables the water depth recognition model to more accurately identify the corresponding water depth information based on the social media content to be identified in the target area and target time period. Ultimately, it can obtain the water depth recognition result of the target area in the target time period based on multiple pieces of social media content to be identified in the target area and target time period.

[0064] In the water depth extraction method based on a direct preference-optimized multimodal large language model of the present invention, before the step of inputting the social media content to be identified into the pre-trained water depth recognition model, the method further includes:

[0065] The multimodal large language model was trained using a direct preference optimization strategy. During the training process, weights for different text modules were added to the original scoring calculation of the direct preference optimization strategy.

[0066] To train a depth recognition model from an initial multimodal large language model, a token-based direct preference optimization strategy is used to fine-tune the initial multimodal large language model. The loss function calculation method based on tokens selectively calculates the loss values ​​between positive and negative samples. The method of directly optimizing for human preferences is based on the answers of the initial multimodal large language model. By comparing the positions of the corrected text, the corresponding text positions are optimized and trained.

[0067] Specifically, direct preference optimization transforms the reinforcement learning objective of human feedback into an equivalent supervised learning objective in a simple way. The key lies in adapting the reward function r(x,y) to the model π to be optimized. * (y|x) and reference model π ref Using (y|x) to represent this allows for direct optimization of the model to be optimized in an appropriate manner based on the preferred or corrected data.

[0068] The reward model r(x,y) can be represented as:

[0069]

[0070] In the formula, β is a constant, and Z(x) is the partition function. Based on this, the model that directly optimizes preferences manually can be expressed as:

[0071]

[0072] The reference model typically consists of a base model that is adjusted by instructions that want to be improved, and is fixed during preference optimization, with only the model to be optimized being updated.

[0073] Furthermore, such as Figure 2 As shown, based on the direct preference optimization strategy, in order to perform more refined feedback and optimization at the text block or token level, it is necessary to add weights to each text module to the original overall ranking scoring method. The original scoring method can be expressed as:

[0074]

[0075] In the formula, y i It is the i-th token in response to y, in order to make the corrected text fragment y c To contribute more to the scoring of the entire text, the above formula can be further modified into a fine-grained weighted aggregation of multiple text fragments:

[0076]

[0077] In the formula, γ is the weight hyperparameter of the corrected segment, γ > 1, and the larger the value of γ, the greater the contribution of the corrected segment to the overall score. N = |y u |+γ|y c |, N is the standardization factor used for normalization to prevent longer answers from receiving higher scores, y u This is the text segment corresponding to the original answer.

[0078] In simple terms, during the training process, the model improves the parts of the multimodal large language model that were wrong in the previous training round, so that it can learn and train more, thereby improving the model's ability to recognize the wrong parts of the answer, and finally training to obtain the water depth recognition model.

[0079] Building on this, in a feasible implementation, the parameters of the multimodal large language model are fine-tuned using a method based on LORA (Low-Rank Adaptation of Large Language Models).

[0080] The principle is to use the weight matrix of the model parameters. By using low-rank decomposition to represent the parameter update ΔW as two smaller low-rank matrices A and B, the number of parameters during model training is reduced.

[0081]

[0082] During training, with parameters W0 frozen and only the parameters of A and B trained, the forward propagation process of a certain layer of the model can be represented as:

[0083] h = W0x + ΔWx = W0x + BAx;

[0084] Fine-tuning of the LORA method model can be viewed as an incremental process on the original parameters. This incremental parameter, ΔW, can be approximated by two low-rank matrices, thus reducing the number of parameters and improving training efficiency.

[0085] Furthermore, in order to evaluate the training effect of the multimodal large language model, in a feasible implementation, the LORA training parameters are first merged, then the test set inference results of the multimodal large language model on different rounds of parameters are merged, and finally, the inference results of each sample on different rounds are sorted according to the quality and correctness of the answer results. The best round is determined according to the average ranking and variance of all samples, the optimal model is selected, and it is determined as the final water depth recognition model.

[0086] After obtaining the water depth identification model through the above method, as follows: Figure 3 As shown, the collected social media content to be identified is input into the model, which processes the data of each modality separately and then inputs it into the large language model to obtain the water depth recognition result output by the large language model. The water depth recognition result describes the water depth of the target area.

[0087] In the water depth extraction method for a multimodal large language model based on direct preference optimization of the present invention, before the step of training the multimodal large language model using the direct preference optimization strategy, the method further includes:

[0088] Obtain multiple social media data related to water depth and filter out social media data that includes geographical location;

[0089] To train a multimodal large language model using a token-based direct preference optimization strategy, a corresponding sample database needs to be constructed.

[0090] Urban water depth extraction first requires acquiring and locating the location information in the text of social media data, which is fundamental for flood monitoring and emergency response. However, the lack of geographic location information in crowdsourced data hinders further spatial analysis and research on flooding events. Currently, only 1%-2% of social media data contains geographic location tags, and these locations do not fully represent the locations described in the text. Most geographic location information is implicit in the text data. Extracting and locating geographic information from crowdsourced data remains challenging. Furthermore, the diverse, flexible, and colloquial nature of textual expressions in social media differs significantly from standard address expressions. Therefore, further exploration is needed.

[0091] In this embodiment, social media data from multiple historical flood disasters is first collected, and data entries containing water depth information and geographical location are selected from them.

[0092] In one specific implementation, social media data for the flood disaster period is collected from predetermined social media platforms, preprocessed, and duplicate text content is removed.

[0093] Optionally, the vector embeddings of all social media texts are first extracted using a text vector model. Then, each piece of data is traversed, and the cosine distance between the current text and all previous texts is calculated. If the distance is less than a certain threshold, such as less than 0.8, the text is considered to be a duplicate of the previous text and is discarded, thus achieving text filtering.

[0094] Based on this, a keyword list is created to remove data from the dataset containing that keyword. The keyword list represents keyword content that is unrelated to flooding or waterlogging but appears frequently.

[0095] Optionally, the keyword list may include the names of online KOLs (Key Opinion Leaders) and celebrities. The specific list can be determined based on the actual data available.

[0096] At this point, we have obtained social media data roughly related to the flood disaster. Then, we use a named entity recognition tool to identify locations within the text, filtering out data without location entities or data lacking coordinate information in the metadata. Understandably, the named entity recognition tool can identify text data in social media data to obtain locations directly mentioned in the text; the metadata is used to obtain location data from the social media data to determine the location based on the location.

[0097] Furthermore, in order to filter out social media data related to flood disasters, some random samples were manually labeled to obtain positive and negative examples of data entries related to flood disasters. The binary classification flood disaster recognition model trained with the BERT model was used to identify positive examples on the dataset after the initial screening.

[0098] Social media data of different modalities are labeled to obtain a dataset used to train the water depth recognition model.

[0099] After collecting social media data related to flood disasters, it is necessary to design targeted processing and labeling for different modalities of social media data types, that is, to determine the labels of the samples and obtain a sample dataset that can be used to train a multimodal large language model.

[0100] In the water depth extraction method based on direct preference optimization of a multimodal large language model of this invention, the step of annotating social media data of different modalities specifically includes:

[0101] For text modal data, extract and standardize the geographical location and water depth information from the text data according to a preset format;

[0102] In one feasible implementation, for text data, the processing and annotation method is to define two named entities: one for geographic location and the other for water depth information.

[0103] Furthermore, considering the diverse ways social media datasets are described, if entities are discontinuous, they are linked together by connecting them from smallest to largest. For example, if the text contains discontinuous expressions such as "Hongshan District: Luoyu Road, Guangba Road", it contains three geographical entities but actually has two geographical locations. In this case, "Luoyu Road" is pointed to "Hongshan District", and "Guangba Road" is pointed to "Hongshan District".

[0104] Furthermore, considering that there may be multiple locations for water depth description in the text, the water depth is associated with the location through a relational link. Finally, a list is generated based on the annotation results. Each element in the list consists of a geographical location and water depth information. If any information is missing, the list is empty.

[0105] For example, [{"Geographical Location":"Luoyu Road, Hongshan District","Water Depth Information":"5cm"}...{...}...]. When input into the model, a set of prompts instructs the model to extract this information in a specific format, such as JSON.

[0106] Among them, the geographical location is the extracted location information, and the water depth information can be direct numerical information such as "5cm" or a description of the water depth, such as "the water is above the instep".

[0107] For image modality data, multiple instructions for querying the image data are predetermined. The instructions and each piece of image data are sequentially input into the multimodal large language model to obtain the recognition result output by the multimodal large language model. Based on the recognition result, the error parts are corrected to obtain the correct response result. The correct response result is used as a positive sample, and the recognition result is used as a negative sample.

[0108] In one feasible implementation, the image data is processed and labeled by first formulating a series of instructions to query the image, which are related to flood disasters and water depth issues, and then inputting the image and instructions into the model to obtain the corresponding water depth information recognition results.

[0109] Based on the recognition results, the incorrect text in the model's response is corrected to obtain the correct response. The corrected text is then used as a positive sample, while the original text of the model's response is used as a negative sample for training. It's understood that the "model" here refers to the initial multimodal large language model.

[0110] For video modal data, a frame of video data is extracted at preset intervals and labeled as image data.

[0111] In one feasible implementation, for video data, the processing and annotation method is to extract a frame of video data at preset intervals as image data, and process and annotate it in the same way as image data.

[0112] Optionally, the preset duration can be determined based on the total duration of the video; in this embodiment, it is set to 1 second.

[0113] Furthermore, it is understandable that when the image is a moving image, it can be labeled in the same way as video data.

[0114] In summary, a multimodal water depth sample database was ultimately formed through manual annotation, which was used to train the multimodal large language model. Specifically, during the training process, the data in the sample database was randomly divided into training and validation sets in an 8:2 ratio.

[0115] In the water depth extraction method based on a direct preference-optimized multimodal large language model of the present invention, after the step of obtaining the water depth recognition result output by the water depth recognition model, the method further includes:

[0116] Predefine water depth levels and classify water depth descriptions, and train a water depth level model based on the defined water depth levels;

[0117] Understandably, since the water depth recognition model is a large language model, the water depth recognition result output by the water depth recognition model is a descriptive text of the input content. In order to more intuitively determine the water depth corresponding to the output descriptive text, a water depth level model can be pre-built and trained after the large language model.

[0118] In one specific implementation, water depth classes are first defined and water depth descriptions are categorized.

[0119] Specifically, such as Figure 4 As shown, based on reference objects in the water depth sample database, including but not limited to people, cars, bicycles, motorcycles, trucks, and buses, different levels are classified according to their described characteristics. The original reference object is people, and the instep, ankle, calf, knee, thigh, hip, waist, chest, neck, and eyes of people in the absence of standing water are divided into 11 levels. Then, the corresponding descriptions of other reference objects are summarized into this category; for example, the middle of a car wheel corresponds to the calf level, and the motorcycle seat corresponds to the thigh level.

[0120] Based on this, a water depth classification model is constructed. Optionally, a pre-trained language model based on BERT (Bidirectional Encoder Representations from Transformers, deep bidirectional language representation model) is used as the base model, followed by a fully connected layer to construct a water depth classification quantization model.

[0121] In other words, a BERT-based deep learning model is used to quantify and classify water depth descriptions, ultimately achieving automatic grading of water depth identification results and determining the water depth identification results as the corresponding water depth level.

[0122] Furthermore, a dataset is generated to train the water depth classification model. Based on the predefined water depth classification, the dataset is randomly divided into a training set and a validation set in an 8:2 ratio. The water depth classification model is trained and evaluated using the dataset, and the best-performing round result on the validation set is selected as the final water depth classification model.

[0123] The water depth identification result is input into the water depth classification model to obtain the water depth classification result output by the water depth classification model.

[0124] After the water depth classification model is trained, it is connected to the water depth identification model. This allows the model to directly output the predefined water depth classification based on the water depth identification results generated by the water depth identification model, thereby obtaining a more intuitive result for identifying the depth of accumulated water.

[0125] In the water depth extraction method based on a direct preference-optimized multimodal large language model of the present invention, after the step of obtaining the water depth level results output by the water depth level model, the method further includes:

[0126] The water depth classification results are then visualized.

[0127] Since the water depth extraction method provided by this invention is mainly used to extract the water depth of water accumulated in urban flooding, the obtained water depth level results can be visualized to more intuitively reflect the water depth situation in different areas of the city.

[0128] In one specific implementation, the social media data is first geolocated using a geocoding API (Application Programming Interface). The returned results then provide location information at different scales.

[0129] Secondly, social media content from the flood disaster period is processed to extract and quantify its water depth information. If a social media data point contains multiple instances of water depth at the same location, aggregation is performed using a truncated averaging method. For example, at the street scale, some data only describe the water depth of the city; these are ignored, and only data containing street locations and more detailed geographic descriptions are considered, from which a truncated average is performed. The method for truncating the average is to truncate the results of multiple instances according to a proportion and then take the average, which can be expressed as:

[0130]

[0131] In the formula, X is the sorted water depth value, n is the number of instances, a is the truncation ratio, g is the integer part of n*a, and r is the fractional part of n*a, which can reduce the influence of extreme values ​​on the final result.

[0132] Based on this, spatiotemporal distribution maps of water depth information at different scales are drawn according to the water depth classification results and geographical location information, completing the visualization processing of the water depth classification results. Thus, the complete process for identifying water depth is as follows: Figure 5 As shown.

[0133] The following describes the water depth extraction device based on the direct preference optimization multimodal large language model provided by the present invention. The water depth extraction device based on the direct preference optimization multimodal large language model described below can be referred to in correspondence with the water depth extraction method based on the direct preference optimization multimodal large language model described above.

[0134] like Figure 6 As shown, the water depth extraction device based on a direct preference-optimized multimodal large language model includes an acquisition module 601 and a recognition module 602.

[0135] The acquisition module 601 is used to acquire the social media content to be identified;

[0136] The social media content to be identified includes text, image, and video modalities, or combinations thereof.

[0137] Specifically, when it is necessary to identify the water depth of a target area at a target time, social media posts containing geographic location information and water depth information are obtained from different social media platforms as the social media content to be identified, thereby identifying the water depth information of the corresponding area based on the social media content to be identified.

[0138] Alternatively, social media can be a social media platform where text, images, or videos are the primary content posted.

[0139] Optionally, text content that directly contains geolocation information can be identified as social media content containing geolocation information, or social media content with location or tags can be identified as social media content containing geolocation information.

[0140] For example, if the text of the social media content to be identified is "There is standing water on Street A, it's up to my calves", it is considered to contain the geographical location information "Street A"; if the social media content to be identified contains the location information "Community B", it is considered to contain the geographical location information "Community B"; if the social media content to be identified contains the hashtag "#CDistrictNews", it is considered to contain the geographical location information "CDistrict C".

[0141] Optionally, if the text, images, and / or videos of social media content contain water depth information, then such content is identified as social media content containing water depth information to be identified.

[0142] Optionally, the image content in social media content includes both still images and moving images.

[0143] Optionally, the social media content to be identified can be one or more pieces.

[0144] The recognition module 602 is used to input the social media content to be recognized into a pre-trained depth recognition model and obtain the depth recognition result output by the depth recognition model, wherein the depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0145] Currently, most methods for extracting depth information based on social media data tend to develop single-modal models for a single data source (such as text or images). This leads to a waste of other modal data resources in social media data and also easily causes bias towards a single modality, ignoring important information carried in other modalities.

[0146] Therefore, in order to achieve water depth identification based on social media data more flexibly, so as to maximize the use of social media data to extract water depth information in a fine-grained manner, this embodiment uses a multimodal large language model as the basis for training to obtain a water depth identification model.

[0147] To fully utilize the diverse modalities of social media data, the selected multimodal large language model needs to be capable of handling large language model data across text, image, and video modalities. This implementation uses MiniCPM-V (an edge-side multimodal large language model) as the base model for illustration.

[0148] MiniCPM-V consists of a visual encoder, a compression layer, and a large language model. It is pre-trained on a Chinese corpus, enabling it to better adapt to social media content in the Chinese context. The visual encoder is responsible for encoding images and video frames, the compression layer is responsible for compressing the visually encoded tokens to a reasonable range, and the large language model is responsible for understanding and outputting the compressed visual and text encodings.

[0149] Although multimodal large language models provide a foundation for unified processing and encoding of data from different modalities, enabling unified understanding and processing by encoding and converting multimodal data into natural language, current large language models generally suffer from illusion problems, lacking sufficient understanding of water depth in images and easily outputting incorrect recognition results. Therefore, in order to unleash the model's ability to understand flood disasters and water levels and improve the model's accuracy, this implementation uses a token-level direct preference optimization strategy to train the selected multimodal large language model to obtain a water depth recognition model with more accurate recognition results.

[0150] The token-level direct preference optimization strategy differs from common direct preference optimization strategies. Instead, it refines the optimization process to the token level. Specifically, during the production response process, text is generated token by token, and the weights of the text segments corresponding to each token are adjusted to achieve more granular adjustment and optimization. This avoids overfitting of the model while fine-tuning the large language model.

[0151] Based on this, the social media content to be identified is input into the trained water depth recognition model, and the water depth recognition result output by the model can be obtained.

[0152] For example, if the text in the social media content to be identified is "The water has flooded my house" and contains the location "D community", and the image content is "The water has submerged the road shoulder", then the water depth recognition model will generate a recognition result such as "The water depth in D community has submerged the road shoulder".

[0153] Understandably, by collecting multiple pieces of social media content to be identified in the target area during the target time period and obtaining their water depth identification results, the water depth information of the target area during the target time period can be obtained based on the water depth identification results.

[0154] This invention trains a multimodal large language model to obtain a water depth recognition model through a token-level direct preference optimization strategy. This enables the water depth recognition model to more accurately identify the corresponding water depth information based on the social media content to be identified in the target area and target time period. Ultimately, it can obtain the water depth recognition result of the target area in the target time period based on multiple pieces of social media content to be identified in the target area and target time period.

[0155] The water depth extraction device based on direct preference optimization multimodal large language model of the present invention also includes a classification module, which is used to predefine water depth levels and classify water depth descriptions, and train a water depth level model based on the defined water depth levels.

[0156] The water depth identification result is input into the water depth classification model to obtain the water depth classification result output by the water depth classification model.

[0157] Understandably, since the water depth recognition model is a large language model, the water depth recognition result output by the water depth recognition model is a descriptive text of the input content. In order to more intuitively determine the water depth corresponding to the output descriptive text, a water depth level model can be pre-built and trained after the large language model.

[0158] In one specific implementation, water depth classes are first defined and water depth descriptions are categorized.

[0159] Specifically, such as Figure 4As shown, based on reference objects in the water depth sample database, including but not limited to people, cars, bicycles, motorcycles, trucks, and buses, different levels are classified according to their described characteristics. The original reference object is people, and the instep, ankle, calf, knee, thigh, hip, waist, chest, neck, and eyes of people in the absence of standing water are divided into 11 levels. Then, the corresponding descriptions of other reference objects are summarized into this category; for example, the middle of a car wheel corresponds to the calf level, and the motorcycle seat corresponds to the thigh level.

[0160] Based on this, a water depth classification model is constructed. Optionally, a pre-trained language model based on BERT (Bidirectional Encoder Representations from Transformers, deep bidirectional language representation model) is used as the base model, followed by a fully connected layer to construct a water depth classification quantization model.

[0161] In other words, a BERT-based deep learning model is used to quantify and classify water depth descriptions, ultimately achieving automatic grading of water depth identification results and determining the water depth identification results as the corresponding water depth level.

[0162] Furthermore, a dataset is generated to train the water depth classification model. Based on the predefined water depth classification, the dataset is randomly divided into a training set and a validation set in an 8:2 ratio. The water depth classification model is trained and evaluated using the dataset, and the best-performing round result on the validation set is selected as the final water depth classification model.

[0163] After the water depth classification model is trained, it is connected to the water depth identification model. This allows the model to directly output the predefined water depth classification based on the water depth identification results generated by the water depth identification model, thereby obtaining a more intuitive result for identifying the depth of accumulated water.

[0164] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a water depth extraction method based on a direct preference optimization multimodal large language model. This method includes: acquiring social media content to be identified; inputting the social media content to be identified into a pre-trained water depth recognition model to obtain the water depth recognition result output by the water depth recognition model, wherein the water depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0165] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0166] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the water depth extraction method based on a direct preference optimization multimodal large language model provided by the above methods. The method includes: acquiring social media content to be identified; inputting the social media content to be identified into a pre-trained water depth recognition model to obtain the water depth recognition result output by the water depth recognition model, wherein the water depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0167] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the water depth extraction method based on a direct preference optimization multimodal large language model provided by the above methods. The method includes: acquiring social media content to be identified; inputting the social media content to be identified into a pre-trained water depth recognition model to obtain the water depth recognition result output by the water depth recognition model, wherein the water depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for extracting water depth based on a direct preference-optimized multimodal large language model, characterized in that, include: Obtain the social media content to be identified; The social media content to be identified is input into a pre-trained depth recognition model to obtain the depth recognition result output by the depth recognition model. The depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy. Specifically, a direct preference optimization strategy was used to train the multimodal large language model. During the training process, weights for different text modules were added to the original scoring calculation of the direct preference optimization strategy. ; In the formula, To correct the weight hyperparameters of the segment, A value greater than 1 indicates that the corrected segment contributes more to the overall score. , N This is a standardization factor used for normalization to prevent longer answers from receiving higher scores. This is the text segment corresponding to the original answer. Is a response y The i One token, The corrected text fragment. For input x Generate response y The scoring function; Before training a multimodal large language model using a direct preference optimization strategy, the following steps are also included: Obtain multiple social media data related to water depth and filter out social media data that includes geographical location; For text modal data, extract and standardize the geographical location and water depth information from the text data according to a preset format; For image modality data, multiple instructions for querying the image data are predetermined. The instructions and each piece of image data are sequentially input into the multimodal large language model to obtain the recognition result output by the multimodal large language model. Based on the recognition result, the error parts are corrected to obtain the correct response result. The correct response result is used as a positive sample, and the recognition result is used as a negative sample. For video modal data, one frame of video data is extracted at preset intervals and labeled as image data. A dataset is obtained for training the water depth identification model.

2. The method for extracting water depth based on a direct preference-optimized multimodal large language model according to claim 1, characterized in that, After the step of obtaining the water depth identification result output by the water depth identification model, the method further includes: Predefine water depth levels and classify water depth descriptions, and train a water depth level model based on the defined water depth levels; The water depth identification result is input into the water depth classification model to obtain the water depth classification result output by the water depth classification model.

3. The method for extracting water depth based on a direct preference-optimized multimodal large language model according to claim 2, characterized in that, After the step of obtaining the water depth level results output by the water depth level model, the method further includes: The water depth classification results are then visualized.

4. A device for extracting water depth based on a direct preference-optimized multimodal large language model, characterized in that, A method for extracting water depth based on a direct preference-optimized multimodal large language model as described in any one of claims 1-3, comprising: The acquisition module is used to acquire the social media content to be identified; The recognition module is used to input the social media content to be recognized into a pre-trained depth recognition model and obtain the depth recognition result output by the depth recognition model, wherein the depth recognition model is a multimodal large language model trained based on a token-level direct preference optimization strategy.

5. The water depth extraction device based on a direct preference-optimized multimodal large language model according to claim 4, characterized in that, Also includes: The grading module is used to predefine water depth grades and classify water depth descriptions, and to train a water depth grade model based on the defined water depth grades. The water depth identification result is input into the water depth classification model to obtain the water depth classification result output by the water depth classification model.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the water depth extraction method based on direct preference optimization of a multimodal large language model as described in any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the water depth extraction method based on direct preference optimization of a multimodal large language model as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method for acquiring flood disaster information from social media

    CN114708485A

  • Urban flood early warning system and method

    CN118840844A