A multi-modal large model-based construction engineering quality evaluation auxiliary system

By extracting and optimizing user interaction behaviors through data collection, corpus analysis, and corpus processing modules, the semantic variation problem of large models in the field of construction engineering is solved, and the accuracy and reliability of information interaction are improved.

CN119740925BActive Publication Date: 2025-11-07GUANGDONG HUAGONG ENG CONSTR SUPERVISION CO LTD +3
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411837000.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-11-07
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

When interacting with large models, especially with special information such as inverted information and misspelled information, existing technologies suffer from semantic ambiguity, leading to misunderstandings and affecting the accuracy and reliability of message responses.

Method used

The data acquisition module obtains historical interaction data between the user terminal and the large model, the corpus analysis module extracts interaction behavior features, the anomaly identification module labels the part-of-speech categories of segmented words, and the corpus processing module optimizes the currently sent information and generates replacement segments to correct semantic anomalies, ensuring that the large model correctly understands the information.

Benefits of technology

It improves the accuracy of information understanding and the reliability of message response in the field of architectural engineering by processing semantic variability, and ensures the accuracy and reliability of information interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740925B_ABST
    Figure CN119740925B_ABST
Patent Text Reader

Abstract

The present application relates to the field of model analysis, and more particularly to a building engineering quality evaluation auxiliary system based on a multi-modal large model, which is provided with a data acquisition module, a corpus analysis module, an anomaly recognition module, a corpus processing module and a large model module, the anomaly recognition module analyzes historical sending information, divides the semantic anomaly of the user end, the corpus analysis module performs part-of-speech category labeling on the historical sending information with strong semantic anomaly, determines the corresponding anomaly tendency part-of-speech category, the corpus processing module processes the sending information based on the semantic anomaly of the user end, so that the sending information can be correctly understood by the large model, the large model module accepts the optimized sending information, and outputs the output result of the large model. The present application constructs a large model dedicated to building engineering quality and safety evaluation, which can more accurately and quickly understand the information in the field of building engineering, and improve the accuracy of retrieval and the reliability of message reply.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of model analysis, in particular to a construction engineering quality evaluation auxiliary system based on a multi-modal large model. BACKGROUND

[0002] Engineering quality problems run through the entire life cycle of engineering construction, and any problem at any link will affect engineering quality and safety, and even cause serious engineering accidents. Especially because of non-information structuring and information fragmentation, there is a large difference in quality and safety management effect. At this time, the informatization construction of engineering quality and safety evaluation is urgent, and with the rapid development of artificial intelligence semantic analysis large model (ChatGPT-like), the networked knowledge system and updating mechanism will serve as an industry standard information artificial intelligence brain. Engineers only need to interact with the decision system, and the system will visually display the best solution and search results to the engineers.

[0003] Chinese patent publication No. CN118735732A discloses an intelligent construction safety management system and method based on a large multi-modal language model. The system includes a data acquisition module, a scene recognition module, a safety behavior analysis module, a warning broadcast module, and a man-machine interaction module. The mobile inspection device and the fixed monitoring device of the data acquisition module are used to collect image data of the construction site in real time, and upload the data to the edge computing node or the cloud server. The pre-trained large multi-modal language model is introduced through the scene recognition module to understand the semantics and classify the scenes of the collected image data, and to realize accurate recognition by combining the scene information library. At the same time, an active learning mechanism is introduced to continuously optimize the model performance. The introduction of the large multi-modal language model significantly improves the understanding and analysis ability of complex construction scenes, realizes intelligent identification and early warning of safety hazards, and effectively reduces the use threshold and improves the management efficiency.

[0004] However, in the prior art, the following problems exist,

[0005] When interacting with a large model, the sent information may have unclear semantics, especially for special information such as inverted information and misspelled information. At this time, the information received by the large model may be incorrect, and the understanding of the information may be biased, resulting in incorrect search direction of the large model, inaccurate content, and generation of incorrect or fabricated information, affecting the reliability of message reply. SUMMARY

[0006] To this end, the application provides a building engineering quality evaluation auxiliary system based on a multi-modal large model, to solve the problem of unclear semantics when interacting with a large model, especially for special information such as inverted information and misspelled information, at this time the information received by the large model will be wrong, the understanding of the information will be biased, the large model will make mistakes in retrieval direction, the content will be inaccurate, and incorrect or fabricated information will be generated, affecting the reliability of message reply.

[0007] To achieve the above-mentioned purpose, the application provides a building engineering quality evaluation auxiliary system based on a multi-modal large model, which comprises:

[0008] A data acquisition module is used to obtain a plurality of historical interaction data of a user terminal and a large model, including historical sending information of the user terminal and historical feedback information of the large model;

[0009] A corpus analysis module is connected with the data acquisition module, used to extract interaction behavior features based on the historical interaction data, to determine semantic anomaly representation parameters of the user terminal based on the interaction behavior features, and to divide the semantic anomaly of the user terminal;

[0010] An anomaly recognition module is connected with the data acquisition module and the corpus analysis module, which is used to mark the part-of-speech category of the segmented segment after processing the historical sending information of the strong semantic anomaly user terminal, and to determine the corresponding anomaly tendency part-of-speech category according to the semantic collocation degree between the segmented segments in each historical sending information;

[0011] A corpus processing module is connected with the anomaly recognition module, used to receive the current sending information of the user terminal, process the sending information based on the semantic anomaly of the user terminal, including,

[0012] Marking the part-of-speech category of each segmented segment in the current sending information, generating a plurality of replacement segments to replace the segmented segments corresponding to the anomaly tendency part-of-speech category, optimizing the sending information according to the replacement result, and recording the optimized current sending information;

[0013] Or, record the current sending information;

[0014] A large model module is connected with the corpus processing module, used to receive the current sending information recorded by the corpus processing module and input to a pre-configured large model, and output the output result of the large model.

[0015] Further, the process of the corpus analysis module to extract interaction behavior features based on the historical interaction data comprises,

[0016] To determine the semantic collocation degree of each historical sending information and record the average semantic collocation degree;

[0017] determine a repeated search behavior based on each historical sending information to determine a repeated search rate;

[0018] The repeated search behavior includes that a semantic collocation degree between the sending information sent by the user terminal in succession is greater than a predetermined semantic collocation degree threshold, and the repeated search rate is a proportion of the sending information corresponding to the repeated search behavior in a total amount of sending information.

[0019] Further, the process in which the corpus analysis module determines the semantic anomaly representation parameter of the user terminal based on the interaction behavior feature includes,

[0020] determining a ratio of the semantic collocation degree threshold to the semantic collocation degree average as a semantic collocation influence factor;

[0021] determining a ratio of the repeated search rate to a repeated search rate threshold as a repeated search influence factor;

[0022] determining a weighted sum of the semantic collocation influence factor and the repeated search influence factor as the semantic anomaly representation parameter.

[0023] Further, the corpus analysis module is configured to divide the semantic anomaly of the user terminal, wherein,

[0024] if the anomaly representation parameter is greater than or equal to a reference anomaly representation parameter, the semantic anomaly is divided into strong semantic anomaly;

[0025] if the anomaly representation parameter is less than the reference anomaly representation parameter, the semantic anomaly is divided into weak semantic anomaly.

[0026] Further, the process in which the anomaly recognition module labels the part-of-speech category of the segmented sentence includes,

[0027] segmenting each historical sending information to obtain a segmented sentence;

[0028] labeling the part-of-speech category of each segmented sentence.

[0029] Further, the process in which the anomaly recognition module determines the anomaly tendency part-of-speech category according to the semantic collocation degree between the segmented sentences in each historical sending information includes,

[0030] determining the semantic collocation degree between the segmented sentence and the remaining segmented sentences in each historical sending information;

[0031] recognizing the anomaly segmented sentence in each historical sending information based on the semantic collocation degree, and labeling the part-of-speech category corresponding to the anomaly segmented sentence;

[0032] determine the probability of the tagged part-of-speech category appearing in the historical sent information, and determine the part-of-speech category with an appearance probability greater than a predetermined probability threshold as the abnormality-prone part-of-speech category;

[0033] If the semantic collocation degree between the segmented sections is less than a predetermined segmented semantic collocation degree threshold, the segmented sections are determined as abnormal segmented sections.

[0034] Further, the corpus processing module is configured to process the sent information based on the semantic abnormality of the user terminal, wherein,

[0035] If the semantic abnormality is strong semantic abnormality, the part-of-speech category of each segmented section in the current sent information is tagged, a number of replacement sections are generated to replace the segmented sections corresponding to the abnormality-prone part-of-speech category, the sent information is optimized according to the replacement result, and the current sent information after optimization is recorded.

[0036] If the semantic abnormality is weak semantic abnormality, the current sent information is recorded.

[0037] Further, the corpus processing module is configured to determine the segmented section corresponding to the abnormality-prone part-of-speech category, comprising,

[0038] determining whether the part-of-speech category of each segmented section is an abnormality-prone part-of-speech category;

[0039] If the part-of-speech category of the segmented section is an abnormality-prone part-of-speech category, it is determined that the segmented section needs to be replaced.

[0040] Further, the corpus processing module is configured to generate a number of replacement sections to replace the segmented sections corresponding to the abnormality-prone part-of-speech category, comprising,

[0041] determining a number of synonyms based on the segmented section, and taking the synonyms as the replacement sections;

[0042] determining a number of strongly associated words based on the segmented section, and taking the strongly associated words as the replacement sections;

[0043] The strongly associated words need to satisfy that the semantic collocation degree between the words and the segmented section is greater than a predetermined strongly associated threshold.

[0044] Further, the corpus processing module is configured to optimize the sent information according to the replacement result, comprising,

[0045] generating a number of sent information after replacing the segmented sections;

[0046] determining the semantic collocation degree of each sent information after replacement;

[0047] determining the sent information after replacement corresponding to the highest semantic collocation degree as the current sent information after optimization.

[0048] Compared with the prior art, the present application sets up a data collection module, a corpus analysis module, an anomaly recognition module, a corpus processing module and a large model module, the anomaly recognition module analyzes the historical sending information, divides the semantic anomaly of the user end, the corpus analysis module performs part-of-speech category labeling on the historical sending information with strong semantic anomaly, determines the corresponding anomaly tendency part-of-speech category, the corpus processing module processes the sending information based on the semantic anomaly of the user end, so that the sending information can be correctly understood by the large model, and the large model module accepts the optimized sending information and outputs the output result of the large model. The present application is aimed at the interaction problem in the field of construction engineering, optimizes the abnormal information, makes the large model understand the information more accurately and quickly, and improves the accuracy of retrieval and the reliability of message reply.

[0049] Especially, the present application calculates semantic anomaly representation parameters to divide semantic anomaly. When the model performs semantic analysis on the interactive content in the field of construction engineering, due to the difference in individual semantic behavior habits, the information sent by the user end may have abnormal writing habits, such as inverted sentences, missing characters and wrong characters, etc. At this time, the model may not be able to accurately identify the sending information. Based on this, the present application integrates the historical sending information of the client with abnormal writing habits, filters out the abnormal sending information, divides the semantic anomaly of the historical sending information, facilitates subsequent division of the semantic anomaly of the user end, and subsequently performs adaptive processing on the sending information sent by different user ends, so that the large model understands the text more accurately and quickly, and improves the accuracy of retrieval and the reliability of message reply.

[0050] Especially, the present application determines the anomaly tendency part-of-speech category of the historical sending information, accurately finds the sending information with strong semantic anomaly, and in actual situations, the semantic anomaly of the user end often has certain rules, such as different language habits from others, dialect-specific habits, etc. Therefore, the present application considers finding the anomaly tendency part-of-speech category, and then representing the part of the language habit that is prone to anomaly, so as to facilitate subsequent targeted optimization of the sending information, make the large model understand the text more accurately and quickly, and improve the accuracy of retrieval and the reliability of message reply.

[0051] Especially, the present application performs double-line processing on the current sending information with different semantic anomaly. For the current sending information with weak semantic anomaly, it is directly recorded for retrieval and reply. For the strong semantic anomaly category, in actual situations, if the current sending information is still directly retrieved and replied, the reply given by the large model will be inaccurate. Based on this, the present application performs part-of-speech category judgment on the current sending information sent by the user end with strong semantic anomaly to optimize the anomaly tendency part-of-speech category, so that the large model understands the text more accurately and quickly, and improves the accuracy of retrieval and the reliability of message reply. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 A structural schematic diagram of a building engineering quality evaluation auxiliary system based on a multi-modal large model according to an embodiment of the application;

[0053] Figure 2 A logic block diagram for dividing semantic anomalies of the user end according to an embodiment of the application;

[0054] Figure 3 A logic block diagram for determining corresponding anomaly tendency part-of-speech categories according to an embodiment of the application. DETAILED DESCRIPTION

[0055] In order to make the objects and advantages of the present application clearer, the present application will be further described in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0056] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not used to limit the protection scope of the present application.

[0057] It should be noted that, in the description of the present application, unless otherwise explicitly specified and limited, the term "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integral connection; it can be mechanical connection, or electrical connection; it can be direct connection, or indirect connection through an intermediate medium, or internal connection of two elements. Those skilled in the art can understand the specific meaning of the above-mentioned term in the present application according to the specific circumstances.

[0058] Please refer to Figures 1-3 as shown, Figure 1 A structural schematic diagram of a building engineering quality evaluation auxiliary system based on a multi-modal large model according to an embodiment of the application, Figure 2 A logic block diagram for dividing semantic anomalies of the user end according to an embodiment of the application, Figure 3 A logic block diagram for determining corresponding anomaly tendency part-of-speech categories according to an embodiment of the application, the building engineering quality evaluation auxiliary system based on a multi-modal large model according to an embodiment of the application, comprising:

[0059] A data acquisition module, which is used to acquire a plurality of historical interaction data of the user end and the large model, including historical sending information of the user end and historical feedback information of the large model;

[0060] It can be understood that the acquisition of the interaction data needs to acquire authorization of the user end.

[0061] a corpus analysis module connected with the data collection module, configured to extract interaction behavior features based on the historical interaction data, to determine semantic anomaly representation parameters of the user end based on the interaction behavior features, and to divide semantic anomaly of the user end;

[0062] an anomaly recognition module connected with the data collection module and the corpus analysis module, configured to perform word segmentation on historical sending information of a strong semantic anomaly user end, to label the part-of-speech categories of the segmented words, and to determine corresponding anomaly tendency part-of-speech categories according to the semantic collocation degrees between the segmented words in each historical sending information;

[0063] a corpus processing module connected with the anomaly recognition module, configured to receive current sending information of the user end, to process the sending information based on semantic anomaly of the user end, including,

[0064] labeling the part-of-speech categories of each segmented word in the current sending information, generating a plurality of replacement segments to replace the segmented words corresponding to the anomaly tendency part-of-speech categories, optimizing the sending information according to the replacement results, and recording the optimized current sending information;

[0065] or, recording the current sending information;

[0066] a large model module connected with the corpus processing module, configured to input the current sending information recorded by the corpus processing module into a pre-configured large model, and to output the output results of the large model.

[0067] Specifically, the training data of the pre-configured large model needs to be pre-processed, specifically including the following steps,

[0068] S01, data collection: collect building standard files, and label the data. It can be understood that the building standard files include national relevant engineering quality standard files and company self-held standard files. It can be understood that the national relevant engineering quality standard files and the company self-held standard files contain engineering requirement index data, component material information, quality detection standard, construction standard, accident handling standard and other multi-dimensional information, thereby providing comprehensive data support in the field of building engineering.

[0069] S02, data cleaning, cleaning and archiving the collected data. It can be understood that data cleaning is to remove error data and invalid data. The specific way of archiving is not limited, and those skilled in the art can divide it according to the needs for convenient classified storage, for example, for the building industry, the text description of quality standard is classified as a category, and the description of error examples is classified as a category. Of course, other ways can also be included, and those skilled in the art can choose according to the needs, which will not be repeated here.

[0070] S03, data storage, can store the archived data to construct a corpus, which will not be repeated here.

[0071] Specifically, for a pre-configured large model applied to the field of construction engineering, it can answer related professional problems in the field of construction engineering. The training process is not limited, and it can be understood that the existing natural language model can be trained through a pre-constructed corpus dedicated to the construction industry. Of course, the architecture of the natural language model is not limited, and those skilled in the art can choose as needed.

[0072] Specifically, in some possible implementations, the large model adopts a multi-model collaboration mode. It can be understood that different models have their own advantages and disadvantages in different tasks, so the multi-model collaboration mode is preferred, for example, using a management engine to manage and train different models.

[0073] In actual application, when using a multi-model collaboration mode, pre-archived corpora of various types can be used to train models suitable for different tasks, such as models for replying to construction quality-related questions and models for replying to error example-related questions. It is worth noting that the above examples are only for demonstration and the purpose is not to limit the scope of protection. Those skilled in the art can train models for different tasks according to their needs, which will not be repeated here.

[0074] Specifically, for the multi-model collaboration mode, a management engine is used for routing. Routing is dynamically assigning tasks to the most suitable model by analyzing the characteristics of the input, ensuring that each model only handles what it is best at. The routing method is not limited, for example, the data feature classifier can determine the data features of the current sent information, and match the most suitable model according to the data features. Of course, the data feature classifier also needs to be trained to ensure that it can accurately distinguish the current sent information and deliver it to different models for output, which will not be repeated here.

[0075] Specifically, in some possible implementations, when multiple models are used, an image processing model can be introduced to analyze the on-site images and provide risk warnings. For example, an image processing model that can recognize worker violations can be pre-trained. The structure of the image processing model is not limited, for example, a neural network architecture image processing model can be used to achieve the corresponding function through pre-training. By combining image, text and multi-modal information, an AI large model capable of processing multi-dimensional data is constructed, which can accurately reply to text information and analyze the actual situation of the construction site to provide risk warnings and other functions, improving the applicability of the large model in the field of construction engineering.

[0076] Specifically, the process in which the corpus analysis module is configured to extract the interaction behavior features based on the historical interaction data comprises,

[0077] determining a semantic collocation degree threshold value based on the semantic collocation degrees of the historical sending information;

[0078] determining a repeated search behavior based on the historical sending information to determine a repeated search rate;

[0079] The repeated search behavior comprises that the inter-sentence semantic collocation degree between the sending information sent by the user terminal successively is greater than a predetermined inter-sentence semantic collocation degree threshold value, and the repeated search rate is a proportion of the sending information corresponding to the repeated search behavior in the total amount of sending information.

[0080] It can be understood that in actual situations, the sending information is mostly single sentence or single paragraph information, and therefore the semantic collocation degree is for single sending information, and the specific calculation method of the semantic collocation degree is not limited, for example, a pre-trained word vector model (such as Word2Vec, GloVe, FastText) can be used to obtain the vector representation of the keywords in the single sending information, and then the cosine similarity between them is calculated, and the average of the cosine similarity between the keywords contained in the sentence can be determined as the semantic collocation degree, which will not be repeated here.

[0081] It can be understood that the inter-sentence semantic collocation degree between the sending information sent by the user terminal successively is for each time of sending information, and the specific method of the inter-sentence semantic collocation degree is not limited, for example, a pre-trained language model such as a BERT model can be used to obtain the vector representation of the sentence generated by the BERT model through the encoder structure to capture the context information in the text, and then the inter-sentence semantic collocation degree between the sentences can be represented by calculating the cosine similarity between the sentences, of course, other methods can also be used, which will not be repeated here.

[0082] Specifically, it can be understood that in actual situations, if the user's language habit is special, there may be a situation that the input information cannot be effectively answered, and then the problem is changed and input multiple times, therefore, if the cosine similarity is used to represent the inter-sentence semantic collocation degree, the inter-sentence semantic collocation degree threshold value is selected in the interval [0.75, 0.85] to represent the situation that the semantics are relatively close.

[0083] Specifically, the process in which the corpus analysis module is configured to extract the interaction behavior features based on the historical interaction data comprises,

[0084] determining a semantic collocation degree threshold value based on the semantic collocation degrees of the historical sending information;

[0085] determine a ratio of the repeated search rate to the repeated search rate threshold value as a repeated search influence factor;

[0086] determine a weighted sum value of the semantic collocation influence factor and the repeated search influence factor as a semantic anomaly representation parameter.

[0087] Specifically, the weight coefficient of the semantic collocation influence factor is 0.6, and the weight coefficient of the repeated search influence factor is 0.4.

[0088] Specifically, the semantic collocation degree threshold value and the repeated search rate threshold value are both obtained by cosine setting, wherein historical interaction data of a plurality of user terminals are obtained in advance, the semantic collocation degree average value and the repeated search rate of the user terminals are recorded, the average value corresponding to the semantic collocation degree average value of each user terminal is solved, the semantic collocation degree threshold value is set as the product of the average value and a first precision coefficient, the repeated search rate average value is solved, the repeated search rate threshold value is set as the product of the repeated search rate average value and a second precision coefficient, the first precision coefficient is selected in the interval [0.75, 0.85], and the second precision coefficient is selected in the interval (0.85, 0.95].

[0089] Specifically, the semantic anomaly representation parameter is calculated to divide the semantic anomaly, and when the model performs semantic analysis on the interactive content in the field of construction engineering, due to the difference in individual semantic behavior habits, the information sent by the user terminal may have abnormal text habits, for example, inverted sentences, missing words, and wrong words. At this time, the model may not be able to accurately identify the sent information. Based on this, the historical sent information of the client with abnormal text habits is integrated, the abnormal sent information is screened out, the semantic anomaly of the historical sent information is divided, the semantic anomaly of the user terminal is divided in the subsequent division, the sent information sent by different user terminals is adaptively processed, the text understanding of the large model is more accurate and fast, and the accuracy of the retrieval and the reliability of the message reply are improved.

[0090] Specifically, the corpus analysis module is used to divide the semantic anomaly of the user terminal, wherein,

[0091] If the anomaly representation parameter is greater than or equal to the reference anomaly representation parameter, the semantic anomaly is divided into strong semantic anomaly;

[0092] If the anomaly representation parameter is less than the reference anomaly representation parameter, the semantic anomaly is divided into weak semantic anomaly.

[0093] Specifically, the reference anomaly representation parameter is obtained by precalculation, a plurality of anomaly representation parameters of the sent information of the user terminal are obtained, the average value of the anomaly representation parameters is calculated, and the reference anomaly representation parameter is determined as 1.24 times of the average value of the anomaly representation parameters.

[0094] Specifically, the process in which the abnormality recognition module is used to label the part-of-speech categories of the segmented segments includes,

[0095] segmenting each historical sending information to obtain segmented segments;

[0096] labeling the part-of-speech categories of each segmented segment.

[0097] Specifically, the segmentation method is not limited, and a segmentation tool can be used to segment to obtain segmented segments. The segmented segments can be keywords. The tool for labeling the part-of-speech categories is not limited, and any part-of-speech category labeling tool in the prior art can be used by those skilled in the art to label, which will not be described here.

[0098] Specifically, the process in which the abnormality recognition module is used to determine the abnormality-tendency part-of-speech categories according to the semantic collocation degrees between the segmented segments in each historical sending information includes,

[0099] determining the semantic collocation degrees between the segmented segments and the remaining segmented segments in each historical sending information;

[0100] recognizing abnormal segmented segments in each historical sending information based on the semantic collocation degrees, and labeling the part-of-speech categories corresponding to the abnormal segmented segments;

[0101] statistically determining the occurrence probabilities of the labeled part-of-speech categories in the historical sending information, and determining the part-of-speech categories with occurrence probabilities greater than a predetermined probability threshold as abnormality-tendency part-of-speech categories;

[0102] If the semantic collocation degree between the segmented segments is less than a predetermined segmented semantic collocation degree threshold, each segmented segment is determined as an abnormal segmented segment.

[0103] Specifically, it can be understood that for the semantic similarity between two segmented segments, a vector representation of the segmented segments can be obtained, and then the cosine similarity between them is calculated, and the cosine similarity is determined as the semantic similarity between the segmented segments.

[0104] Specifically, the segmented semantic collocation degree threshold is obtained by prior setting. Specifically, the semantic collocation degrees between the segmented segments in a plurality of single sending information are obtained in advance, the average of the semantic collocation degrees is solved, and 0.65 to 0.85 times of the average of the semantic collocation degrees is set as the segmented semantic collocation degree threshold.

[0105] Specifically, the purpose of setting the predetermined probability threshold is to identify part-of-speech categories with high probabilities. Therefore, in order to ensure data representation and exclude occasional situations, the probability threshold needs to be greater than 10% when setting.

[0106] Specifically, the application determines the part-of-speech category of the deviant tendency of the historical sending information segment, accurately searches for the sending information with strong semantic deviation, and in actual situations, the semantic deviation of the user end often has certain rules, for example, the use habit is different from others, and the dialect has certain regularity, therefore, the application considers that the part of speech habit is easy to deviate, and the subsequent sending information is optimized, so that the large model understands the text more accurately and quickly, and the accuracy of the search and the reliability of the message reply are improved.

[0107] Specifically, the corpus processing module is used to process the sending information based on the semantic deviation of the user end, wherein,

[0108] If the semantic deviation is strong semantic deviation, the part-of-speech category of each segmented segment in the current sending information is marked, a plurality of replacement segments are generated to replace the segmented segment corresponding to the deviant tendency part-of-speech category, the sending information is optimized according to the replacement result, and the current sending information after optimization is recorded;

[0109] If the semantic deviation is weak semantic deviation, the current sending information is recorded.

[0110] Specifically, the application processes the current sending information with different semantic deviations in two lines, records the current sending information with weak semantic deviation for retrieval and reply, and determines the strong deviant tendency part-of-speech category of the current sending information with strong semantic deviation for optimization in actual situations. If the current sending information is still directly retrieved and replied, the reply given by the large model will be inaccurate, so the application determines the strong deviant tendency part-of-speech category of the current sending information with strong semantic deviation for optimization, replaces the near-synonymous words or strong associated words, so that the large model understands the text more accurately and quickly, and the accuracy of the search and the reliability of the message reply are improved.

[0111] Specifically, the corpus processing module is used to determine the segmented segment corresponding to the deviant tendency part-of-speech category, comprising,

[0112] determining whether the part-of-speech category of each segmented segment is a deviant tendency part-of-speech category;

[0113] If the part-of-speech category of the segmented segment is a deviant tendency part-of-speech category, it is determined that the segmented segment needs to be replaced.

[0114] Specifically, the process of the corpus processing module for generating a plurality of replacement segments to replace the segmented segment corresponding to the deviant tendency part-of-speech category comprises,

[0115] determining a plurality of synonyms based on the segmented segment, and using the synonyms as replacement segments;

[0116] determine a plurality of strongly associated words based on the segmented word, and use the strongly associated words as replacement words;

[0117] The strongly associated words need to satisfy that the semantic collocation degree between the words and the segmented word is greater than a predetermined strongly associated threshold.

[0118] Specifically, the strongly associated threshold is obtained by pre-setting, wherein the semantic collocation degrees between the segmented words in a plurality of single-time sending information are obtained in advance, the average of the semantic collocation degrees is calculated, and 1.45 to 1.85 times of the average of the semantic collocation degrees is set as the strongly associated threshold. Of course, the strongly associated threshold has an upper limit, which can be set as 0.95. If the calculated strongly associated threshold is the upper limit, the strongly associated threshold is set as the upper limit.

[0119] Specifically, the process of the corpus processing module for optimizing the sending information according to the replacement result includes,

[0120] generating a plurality of sending information after replacing the segmented word;

[0121] determining the semantic collocation degree of each replaced sending information;

[0122] determining the replaced sending information corresponding to the highest semantic collocation degree as the optimized current sending information.

[0123] Specifically, for a sending information, there may be a plurality of alternative words. Replacing the plurality of alternative words in the sending information will generate a plurality of replaced sending information. Determining the sending information with the highest semantic collocation degree as the optimized sending information can enable the large model module to better complete the building engineering quality evaluation auxiliary work.

[0124] The technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings. However, it is easily understood by those skilled in the art that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application. The technical solutions after the changes or replacements will fall within the protection scope of the present application.

Claims

1.A multi-modal large model-based construction quality assessment assistance system, characterized by, Comprise: a data collection module configured to obtain a plurality of historical interaction data between a user terminal and a large model, including historical sending information of the user terminal and historical feedback information of the large model; a corpus analysis module connected with the data collection module, configured to extract interaction behavior features based on the historical interaction data, determine semantic anomaly representation parameters of the user terminal based on the interaction behavior features, and divide semantic anomaly of the user terminal, including: determining the ratio of the semantic collocation degree threshold value to the semantic collocation degree average as a semantic collocation influence factor; determining the ratio of the repeated search rate to the repeated search rate threshold value as a repeated search influence factor; determining the weighted sum of the semantic collocation influence factor and the repeated search influence factor as the semantic anomaly representation parameter; if the anomaly representation parameter is greater than or equal to the reference anomaly representation parameter, dividing the semantic anomaly into strong semantic anomaly; an anomaly recognition module connected with the data collection module and the corpus analysis module, configured to perform part-of-speech tagging on the historical sending information of the strong semantic anomaly user terminal after processing, and determine the part-of-speech category of each part-of-speech segment, including: configured to determine the semantic collocation degree between each part-of-speech segment and the remaining part-of-speech segments in each historical sending information; configured to identify anomaly part-of-speech segments in each historical sending information based on the semantic collocation degree, wherein if the semantic collocation degree between part-of-speech segments is less than a predetermined part-of-speech semantic collocation degree threshold value, each part-of-speech segment is determined to be an anomaly part-of-speech segment, and the part-of-speech category of the anomaly part-of-speech segment is labeled, including: configured to calculate the appearance probability of the labeled part-of-speech category in the historical sending information, and determine the part-of-speech category with an appearance probability greater than a predetermined probability threshold value as an anomaly-tendency part-of-speech category; a corpus processing module connected with the anomaly recognition module, configured to receive current sending information of the user terminal, and process the sending information based on the semantic anomaly of the user terminal, including: if the semantic anomaly is strong semantic anomaly, label the part-of-speech category of each part-of-speech segment in the current sending information, if the part-of-speech category of the part-of-speech segment is an anomaly-tendency part-of-speech category, determine that the part-of-speech segment needs to be replaced, generate a plurality of replacement segments to replace the part-of-speech segment corresponding to the anomaly-tendency part-of-speech category, optimize the sending information according to the replacement result, and record the optimized current sending information; a large model module connected with the corpus processing module, configured to input the current sending information recorded by the corpus processing module into a preconfigured large model, and output the output result of the large model. 2.The multi-modal large model-based construction quality evaluation assistance system of claim 1, wherein, The process of the corpus analysis module for extracting interaction behavior features based on the historical interaction data includes: configured to determine the semantic collocation degree of each historical sending information to record the semantic collocation degree average; configured to determine repeated search behavior based on each historical sending information to determine the repeated search rate; wherein the repeated search behavior includes that the inter-sentence semantic collocation degree between the sending information sent by the user terminal in succession is greater than a predetermined inter-sentence semantic collocation degree threshold value, and the repeated search rate is the proportion of the sending information corresponding to the repeated search behavior in the total amount of sending information. 3.The multi-modal large model-based construction quality evaluation assistance system of claim 1, wherein, The corpus analysis module is used to divide semantic abnormality of the user terminal, wherein, If the abnormality representation parameter is less than the reference abnormality representation parameter, the semantic abnormality is divided into weak semantic abnormality. 4.The multi-modal large model-based construction quality evaluation assistance system of claim 1, wherein, The abnormality recognition module is used to mark the part-of-speech category of the segmented segment, including, segmenting each historical sending information to obtain a segmented segment; marking the part-of-speech category of each segmented segment. 5.The multi-modal large model based construction quality assessment assistant system according to claim 1, wherein, The corpus processing module is used to process the sending information based on the semantic abnormality of the user terminal, wherein, If the semantic abnormality is weak semantic abnormality, the current sending information is recorded. 6.The multi-modal large model-based construction quality evaluation assistance system of claim 1, wherein, The corpus processing module is used to generate a number of replacement segments to replace the segmented segment corresponding to the abnormality tendency part-of-speech category, including, determining a number of synonyms based on the segmented segment, and taking the synonyms as the replacement segment; determining a number of strong correlation vocabularies based on the segmented segment, and taking the strong correlation vocabularies as the replacement segment; Wherein, the strong correlation vocabulary needs to satisfy that the semantic collocation degree between the vocabulary and the segmented segment is greater than a predetermined strong correlation threshold. 7.The multi-modal large model based construction quality assessment assistant system according to claim 1, wherein, The corpus processing module is used to optimize the sending information according to the replacement result, including, generating a number of sending information after replacing the segmented segment; determining the semantic collocation degree of each replaced sending information; determining the replaced sending information corresponding to the highest semantic collocation degree as the optimized current sending information.

Citation Information

Patent Citations

  • Intelligent construction safety management system and method based on large-scale multi-modal language model

    CN118735732A

  • Dynamic data desensitization method

    CN118332604A

  • Green building intelligent evaluation model based on large language model

    CN118863229A