Large model-based data processing method, apparatus and device, and program product
By using a large model to perform semantic similarity matrix analysis on SQL statements and data table fields, the problems of low efficiency and accuracy of manual detection are solved, efficient and accurate field sequence detection is achieved, and data errors and redundancy are avoided.
Patent Information
- Application Number
- CN202511134626.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-09-12
AI Technical Summary
In the prior art, manual detection of SQL statement field order misalignment is inefficient and inaccurate, especially when there are massive SQL statements and target data tables that change frequently. It is difficult to detect field order misalignment problems.
Through the large model, a semantic similarity matrix analysis is performed on the fields in the target data table and SQL statements, the field order is adjusted, and the semantic similarity matrix is used to determine whether the SQL statement field order is normal or misplaced.
It improves the accuracy and efficiency of SQL statement field sequence detection, reduces data errors and redundancy, and ensures data quality.
Smart Images

Figure CN120631996A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of large models and data processing technology, and in particular, to a data processing method, apparatus, device, and program product based on a large model. Background Art
[0002] Field order misalignment refers to data writing or updating errors caused by the mismatch between the field order in the SQL (Structured Query Language) statement and the field order defined in the target data table during database operations.
[0003] In related technologies, field order detection can be performed manually on SQL statements. However, in the face of massive SQL statements and frequent changes in target data tables, manual detection is not only inefficient but also difficult to detect field order misalignment problems in SQL statements, resulting in low detection accuracy. Summary of the Invention
[0004] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides a data processing method based on a large model, the data processing method comprising: Determining a table field in a target data table and a statement field in a target SQL statement, wherein the target SQL statement is used to write data to the target data table or update data in the target data table; When the number of the table fields and the number of the sentence fields are equal, semantic understanding is performed on each pair of the table fields and the sentence fields using the large model to obtain a first semantic similarity matrix; Adjusting a field order of a first field according to the first semantic similarity matrix, and determining a second semantic similarity matrix corresponding to the adjusted first field and a second field, wherein the first field is the table field or the statement field, the second field is the table field or the statement field, and the first field and the second field are different; A recognition result of the target SQL statement is determined according to the first semantic similarity matrix and the second semantic similarity matrix, where the recognition result is used to characterize whether the order of fields in the target SQL statement is normal or abnormal.
[0006] In a second aspect, the present disclosure provides a data processing device based on a large model, the data processing device comprising: a determination module, configured to determine table fields in a target data table and statement fields in a target SQL statement, wherein the target SQL statement is used to write data to the target data table or update data on the target data table; a semantic understanding module, configured to, when the number of the table fields and the number of the sentence fields are equal, perform semantic understanding on each pair of the table fields and the sentence fields using a large model to obtain a first semantic similarity matrix; an adjustment module, configured to adjust a field order of a first field according to the first semantic similarity matrix, and determine a second semantic similarity matrix corresponding to the adjusted first field and a second field, wherein the first field is the table field or the statement field, the second field is the table field or the statement field, and the first field and the second field are different; The recognition module is used to determine a recognition result of the target SQL statement based on the first semantic similarity matrix and the second semantic similarity matrix, wherein the recognition result is used to characterize whether the order of fields in the target SQL statement is normal or abnormal.
[0007] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which implements the steps of the method described in the first aspect when executed by a processing device.
[0008] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of the method in the first aspect.
[0009] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.
[0010] Through the above technical solution, the table fields in the target data table and the statement fields in the target SQL statement are first determined. When the number of table fields and statement fields is equal, the large model is used to perform semantic understanding on each pair of table fields and statement fields, obtaining a first semantic similarity matrix. Then, a second semantic similarity matrix corresponding to the first field and the second field after adjusting the field order is determined. Finally, based on the first semantic similarity matrix and the second semantic similarity matrix, an identification result is determined to indicate whether the target SQL statement field order is normal or abnormal. Using this method, the large model is used to determine the similarity between the table fields and statement fields from a semantic dimension, and the field order is adjusted to further analyze whether the SQL statement has a field order misalignment problem based on the two semantic similarity matrices. This method not only improves detection accuracy but also improves detection efficiency, thereby being applicable to large-scale SQL statement detection scenarios. In addition, in actual business scenarios, the identification results can also avoid data errors caused by executing SQL statements with misaligned field orders, thereby ensuring data quality, reducing data redundancy caused by erroneous data, and saving data storage resources.
[0011] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flow chart illustrating a data processing method based on a large model according to an exemplary embodiment of the present disclosure; Figure 2 is a process diagram of a data processing method based on a large model according to an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram showing a semantic similarity matrix according to an exemplary embodiment of the present disclosure; Figure 4 is a schematic diagram showing a comparison between a statement field and a table field according to an exemplary embodiment of the present disclosure; Figure 5 This is a structural block diagram of a data processing method and apparatus based on a large model according to an exemplary embodiment of the present disclosure; Figure 6 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0013] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0014] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0015] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0017] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0019] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0020] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0021] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0022] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0023] At the same time, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0024] Field order misalignment occurs when data is written or updated incorrectly during database operations due to a mismatch between the field order in the SQL (Structured Query Language) statement and the target table's defined field order. For example, during data writing, data values may be incorrectly mapped to unexpected table columns due to a mismatch between the field order in the SQL write statement and the target table's defined field order. For example, if the target table's field order is a, b, c, and the SQL write statement's field order is c, b, a, the value that should have been written to field a is written to field c, and vice versa. In this case, even if the field names and data types are completely correct, the order discrepancy can still cause data logic errors, compromising data quality.
[0025] There are many reasons for field order misalignment. For example, when writing SQL statements, users may enumerate fields in the wrong order, leading to field order misalignment. Alternatively, the target table's field order may be changed without updating the SQL statements for related data processing tasks. For example, adding a new field to the target table causes subsequent fields to shift position, which can lead to data misalignment and quality issues when executing SQL statements for related data processing tasks. Alternatively, the SQL statement developer and the target table developer may have inconsistently named the same business field, leading to field order misalignment.
[0026] In related technologies, field order detection can be performed manually on SQL statements. However, in the face of massive SQL statements and frequent changes in target tables, manual detection is not only inefficient but also difficult to detect field order misalignment problems in SQL statements, resulting in low detection accuracy.
[0027] Alternatively, you can use a preset regular expression to extract the field names of the SQL statement. However, only the fields explicitly named in the SQL statement can be extracted. If the fields in the SQL statement are not named with the AS keyword or are inconsistent with the field names of the target table, the field order in the SQL statement will be judged to be misplaced, resulting in low detection accuracy.
[0028] Alternatively, a text similarity algorithm is used to calculate the text similarity between the field names in the SQL statement and the field names in the target table, and whether there is a field order misalignment is determined based on whether the field names are consistent. However, this method has detection blind spots. For example, "abcdefg" and "abc" are mapping fields for name abbreviations, "abcdefg" and "xyz" are mapping fields for name variants, and "abcdefg" and "abcefg" are mapping fields after spelling errors. The fields in the above examples are actually equivalent, but they will be judged as inconsistent based on the text similarity algorithm. In addition, fields with consistent semantics but inconsistent names cannot be detected. For example, the semantic equivalence of the two fields "mobile phone number" and "contact phone number" will be judged as inconsistent based on the text similarity algorithm, so the detection accuracy is not high.
[0029] In view of this, the present disclosure provides a data processing method, device, equipment and program product based on a large model to solve the above technical problems.
[0030] The following further explains the embodiments of the present disclosure with reference to the accompanying drawings.
[0031] Figure 1 This is a flow chart of a data processing method based on a large model according to an exemplary embodiment of the present disclosure, referring to Figure 1 , the data processing method may include the following steps: S101: Determine table fields in a target data table and statement fields in a target SQL statement.
[0032] The target SQL statement is used to write data to the target data table or update data in the target data table.
[0033] For example, the target SQL statement may be a DML (Data Manipulation Language) statement, such as an INSERT statement, etc., which is not limited in the present disclosure.
[0034] In a possible manner, determining the table fields in the target data table and the statement fields in the target SQL statement includes: parsing the table structure information of the target data table to obtain the first original field corresponding to the target data table, parsing the target SQL statement to obtain the second original field corresponding to the target SQL statement; determining the target fields having the same arrangement sequence number and the same field name in the first original field and the second original field; excluding the target field from the first original field to obtain the table field, and excluding the target field from the second original field to obtain the statement field.
[0035] For example, Figure 2 As shown, the SQL statement to be identified is obtained, the SQL statement is parsed based on the lexical and grammatical rules of SQL, and the table name of the target data table and the second original field in the SQL statement where data is to be written or updated are extracted, and the second original field is arranged in the order of the fields in the SQL statement, such as the value list in the VALUES clause or the SELECT clause, as well as the processing logic of the statement field, etc., which can be specifically set according to needs and is not limited in this disclosure.
[0036] For example, Figure 2 As shown, based on the table name of the target data table, the interface is called to query the DDL (Data Definition Language) statement of the target data table from the database metadata or version-controlled file used to define the table structure of the database, and the first original field defined by the target data table is parsed therefrom, and the first original field is arranged in the order of the fields defined in the table.
[0037] It should be noted that you can first check whether the number of fields in the first original field and the second original field is consistent. If the number of fields is inconsistent, you can directly determine that there is an anomaly. If the number of fields is consistent, you can identify the field positions in the first original field and the second original field that are in the same position but have different field names, and eliminate the fields with exactly the same field names to obtain table fields and statement fields, thereby obtaining potential misplaced fields, so as to reduce subsequent computing consumption, further save computing resources, and speed up detection efficiency.
[0038] In a possible embodiment, the data processing method further includes: performing at least one of the following preprocessing on the second original field: if a third original field without a specified field name exists in the second original field, completing the field name of the third original field based on the processing logic of the third original field in the target SQL statement; standardizing the field name of the second original field, wherein the standardization includes at least one of unifying uppercase and lowercase letters and removing special characters; and correcting the field name of the second original field that has spelling errors. Determining target fields with the same arrangement number and the same field name in the first original field and the second original field includes: determining target fields with the same arrangement number and the same field name in the first original field and the preprocessed second original field.
[0039] For example, Figure 2 As shown, if there is a field c without a specified field name in the SQL statement, the field name is completed according to the field processing logic. For example, the field value of field c is obtained by adding the field value of field a and the field value of field b, then the field name of field c can be completed as "a_b_total", and so on. Alternatively, you can search in the first original field to find out whether there is a field for storing the total value of the field value of field a and the field value of field b, and use the field name of this field as the field name of field c. The specific field name can be determined according to the actual business scenario, and the present disclosure does not impose any restrictions on this.
[0040] For example, field names can also be standardized, such as unifying upper and lower case, removing special characters such as quotation marks, etc., and obvious field name spelling errors can be checked and corrected based on the edit distance algorithm. Here, it is possible to determine whether the field name is spelled incorrectly based on the word rules, and it can also be detected based on the field name in the first original field. For example, assuming that the first original field includes a field with the field name "abcdefg" and the second original field includes a field with the field name "acdefg", it can be determined that there is a field name spelling error, and then corrections can be made to obtain the correct field name.
[0041] Furthermore, by detecting identical fields between the completed and corrected second original field and the first original field, judgment errors can be reduced. In other words, by completing and correcting field names, the system has good compatibility and adaptability for situations such as missing field names, incorrect field names, and inconsistent field names. It can more accurately find out-of-order fields, further improving the efficiency of detecting whether there are misaligned field sequences.
[0042] S102: When the number of table fields and sentence fields is equal, semantic understanding is performed on each pair of table fields and sentence fields using the large model to obtain a first semantic similarity matrix.
[0043] It should be noted that, in addition to the above-mentioned detection of whether the number of fields is equal in the original field, it is also possible to detect whether the number of fields is equal after determining the table field and the statement field. The specific setting can be based on demand, and this disclosure does not impose any restrictions on this.
[0044] In a possible embodiment, the data processing method further includes: parsing a target SQL statement to obtain first information representing processing logic for fields in the target SQL statement, and parsing table structure information of a target data table to obtain second information, the second information including a table name of the target data table and / or field description information in the target data table. Performing semantic understanding on each pair of table fields and statement fields using a large model to obtain a first semantic similarity matrix includes: performing semantic understanding on each pair of table fields and statement fields using the large model based on the first information, the second information, the field names of the table fields, and the field names of the statement fields to obtain the first semantic similarity matrix.
[0045] For example, the above field pairs include field pairs obtained by combining each table field with each statement field.
[0046] For example, Figure 2 As shown, the big model can perform semantic similarity analysis based on the field names of table fields and statement fields, and can also perform a complete analysis of the statement field in combination with the table name of the target data table, field processing logic and other information, as well as a comprehensive analysis of the table field based on the table name, field description and other information, so that the big model can comprehensively calculate the similarity between the two fields from the aspects of field processing logic, field name, field description, etc. Compared with the big model that simply judges semantic similarity based on field name, the big model in this embodiment can effectively improve the accuracy of judging field similarity by combining more information, thereby improving the accuracy of detecting whether there is a misalignment anomaly in the field order.
[0047] S103: Adjusting the field order of the first field according to the first semantic similarity matrix, and determining a second semantic similarity matrix corresponding to the adjusted first field and the second field.
[0048] The first field is a table field or a statement field, the second field is a table field or a statement field, and the first field is different from the second field.
[0049] S104: Determine a recognition result of the target SQL statement according to the first semantic similarity matrix and the second semantic similarity matrix, where the recognition result is used to indicate whether the order of fields in the target SQL statement is normal or abnormal.
[0050] By adopting the above method, the similarity between table fields and statement fields is determined from the semantic dimension through a large model, and the field order is adjusted to further analyze whether there is a field order misalignment problem in the SQL statement based on the two semantic similarity matrices. This can not only improve the detection accuracy, but also improve the detection efficiency, so that it can be applied to the detection scenario of large-scale SQL statements. In addition, in actual business scenarios, data errors caused by executing SQL statements with misaligned field orders can be avoided based on the recognition results, thereby ensuring data quality, reducing data redundancy caused by erroneous data, and saving data storage resources.
[0051] In a possible manner, the semantic similarity corresponding to the i-th row and j-th column in the first semantic similarity matrix represents the semantic similarity between the i-th field in the statement field and the j-th field in the table field, where i and j are positive integers. According to the first semantic similarity matrix, the field order of the first field is adjusted, including: for each semantic similarity on the main diagonal of the first semantic similarity matrix, when the semantic similarity is the maximum value in the row, determining the semantic similarity as the target semantic similarity; determining the third field corresponding to the target semantic similarity in the first field, and determining a first average value of the remaining semantic similarities on the main diagonal except the target semantic similarity; when the first average value is less than a first preset threshold, adjusting the field order of the fourth field in the first field except the third field.
[0052] For example, since the number of table fields and statement fields is equal, for example, N fields, then the obtained N×N semantic similarity matrix, i and j are positive integers less than or equal to N. Figure 2 As shown, assuming that 1 represents complete similarity and 0 represents complete dissimilarity, the semantic similarity between each pair of table fields and statement fields is determined by the large model, and the following is obtained: Figure 3 The first semantic similarity matrix shown, where the semantic similarity corresponding to the i-th row and j-th column represents the semantic similarity between the i-th field in the statement field and the j-th field in the table field, and the value on the main diagonal is the semantic similarity of the current field order.
[0053] For example, the semantic similarity of the main diagonal of the semantic similarity matrix is traversed. If the semantic similarity is the maximum value in the same row, it means that the corresponding table field and statement field have the same arrangement number and are the most semantically similar fields, and can be removed from the potential misaligned fields, such as Figure 3 The black thick boxes shown in the figure show the table fields and statement fields corresponding to the semantic similarity, thereby further reducing the subsequent computational cost.
[0054] Furthermore, if Figure 2As shown, the average of the remaining semantic similarities on the main diagonal is calculated to obtain the remaining average semantic similarity, and then a determination is made as to whether the remaining average semantic similarity is below a threshold. The first preset threshold can be a fixed value set as required, or can be adjusted based on the number of table fields. For example, the threshold can be appropriately lowered with more fields. A higher threshold can also be used for key tables based on business scenarios, and so on. The specific setting can be based on requirements, and this disclosure does not impose any restrictions on this.
[0055] It should be noted that if Figure 2 As shown in , if the average value of the remaining semantic similarity on the main diagonal is lower than the set threshold, e.g. Figure 3 If the average value of the matrix elements corresponding to [B1, B2] and [C1, C2] is obtained, it can be considered that the corresponding statement fields have the risk of field order misalignment. Then the field order of the corresponding table fields or statement fields can be adjusted to obtain a new semantic similarity matrix for further verification, such as Figure 3 As shown in the figure, assuming the first field is a table field, the fourth field is table fields B2 and C2. The order of table fields B2 and C2 is then adjusted to obtain a new semantic similarity matrix. By first eliminating fields with the same order number and field name, and then eliminating fields with the same order number and the most similar semantics, potential misaligned fields are obtained for further verification. This not only reduces computational overhead but also facilitates the precise location of fields with misaligned order in SQL statements.
[0056] In a possible embodiment, adjusting the field order of the fourth field excluding the third field in the first field includes: when the number of the fourth fields is less than or equal to a second preset threshold, traversing all possible field orders of the fourth field to obtain at least one adjusted first field. Determining a second semantic similarity matrix corresponding to the adjusted first field and the second field includes: determining the semantic similarity matrix corresponding to each adjusted first field and the second field to obtain multiple semantic similarity matrices; and determining, among the multiple semantic similarity matrices, a semantic similarity matrix corresponding to the fourth field on the main diagonal with the maximum average semantic similarity as the second semantic similarity matrix.
[0057] For example, in the process of adjusting the field order, a preset threshold can be set. If the number of the fourth field is less than or equal to the preset threshold, it means that the number of the fourth field is small. All possible field orders can be traversed, and the field order with the largest average semantic similarity corresponding to the fourth field on the main diagonal is taken, and the corresponding semantic similarity matrix is used as the second semantic similarity matrix. This not only facilitates the subsequent verification of whether there is a field order misalignment problem in the SQL statement, but also can analyze the possible correct field order based on this. In other words, not only can the fields with misaligned order be accurately located, but correction suggestions can also be given or the SQL statement can be automatically corrected.
[0058] In a possible embodiment, adjusting the order of the fourth fields in the first field, excluding the third field, includes randomly adjusting the order of the fourth fields in the first field to obtain the fifth field when the number of the fourth fields is greater than a second preset threshold, until the average semantic similarity corresponding to the fourth fields on the main diagonal of the semantic similarity matrix corresponding to the fifth field and the second field is greater than the first average value. Determining the adjusted second semantic similarity matrix corresponding to the first field and the second field includes determining the semantic similarity matrix corresponding to the fifth field and the second field as the second semantic similarity matrix.
[0059] For example, in the process of adjusting the order of fields, a preset threshold can be set. If the number of the fourth field is greater than the preset threshold, it means that the number of the fourth field is large. Continuing with the adjustment of the table field as an example, the field order of the fourth field in the table field can be randomly adjusted to obtain a new semantic similarity matrix, and then the average semantic similarity corresponding to the fourth field on the main diagonal in the new semantic similarity matrix is determined. If the average semantic similarity is greater than the average semantic similarity corresponding to the fourth field on the main diagonal in the semantic similarity matrix corresponding to the original field order, it means that the table field after adjusting the field order is closer to the statement field semantics, and the field order adjustment can be stopped, and the semantic similarity matrix is determined as the second semantic similarity matrix.
[0060] Accordingly, if the average semantic similarity corresponding to the fourth field on the main diagonal in the new semantic similarity matrix is less than or equal to the average semantic similarity corresponding to the fourth field on the main diagonal in the semantic similarity matrix corresponding to the original field order, it means that the semantics of the table field and the statement field are not closer after the field order is adjusted, then the field order of the fourth field in the table field will continue to be randomly adjusted.
[0061] This can minimize the number of adjustments and reduce computing consumption when there are a large number of potential misaligned fields, and also improve the efficiency of detecting SQL statement field order misalignment anomalies.
[0062] It should be noted that the matrix element values in the second semantic similarity matrix can be obtained by adjusting the matrix element values in the first semantic similarity matrix through the adjusted field order, or they can be directly recalculated through the large model. The specific settings can be based on the needs, and this disclosure does not impose any restrictions on this.
[0063] In a possible manner, determining the recognition result of the target SQL statement based on the first semantic similarity matrix and the second semantic similarity matrix includes: determining a second average value of the semantic similarity corresponding to the fourth field on the main diagonal of the first semantic similarity matrix, and determining a third average value of the semantic similarity corresponding to the fourth field on the main diagonal of the second semantic similarity matrix; and determining the recognition result of the target SQL statement based on the second average value and the third average value.
[0064] For example, Figure 3 As shown, the average semantic similarity of the potential misplaced fields on the diagonal of the first semantic similarity matrix and the average semantic similarity of the potential misplaced fields on the diagonal of the second semantic similarity matrix obtained by adjusting the field order are calculated respectively, and then as shown in FIG. Figure 2 As shown in FIG, by comparing the two average semantic similarities to further analyze whether there is a field order misalignment problem in the SQL statement, not only the detection accuracy can be improved, but also the detection efficiency can be improved.
[0065] In a possible manner, based on the second average value and the third average value, determining the recognition result of the target SQL statement, including: when the difference obtained by subtracting the second average value from the third average value is greater than a third preset threshold, determining the recognition result of the field order misalignment abnormality characterizing the target SQL statement; or, when the ratio of the difference obtained by subtracting the second average value from the third average value to the second average value is greater than a fourth preset threshold, determining the recognition result of the field order misalignment abnormality characterizing the target SQL statement.
[0066] For example, Figure 2 As shown, the average semantic similarity of the potential misplaced fields on the diagonal of the adjusted semantic similarity matrix is significantly improved compared with the average semantic similarity of the potential misplaced fields on the diagonal of the original semantic similarity matrix, which can be determined that there is a field order misplacement anomaly in the sentence fields.
[0067] For example, the difference between the average semantic similarity after adjustment and the average semantic similarity before adjustment can be determined. If this difference is greater than a third preset threshold, it indicates that the semantic similarity of the fields after adjustment has significantly improved. Alternatively, the ratio of the difference to the average semantic similarity before adjustment can be further calculated. If this ratio is greater than a fourth preset threshold, it also indicates that the semantic similarity of the fields after adjustment has significantly improved. Thus, through reverse verification, when the semantic similarity of the fields determined by adjusting the field order improves, it is determined that there is a field order misalignment anomaly in the statement field, thereby reducing the detection error rate and further improving the detection efficiency of SQL statement field order misalignment anomalies.
[0068] like Figure 4 As shown in the figure, for the statement fields extracted from the SQL statement and the table fields extracted from the target data table, the fields with the same arrangement position and the same field name are first removed. Then, the semantic understanding of the remaining statement fields and table fields is performed on the field pairs of two combinations through the large model to obtain the semantic similarity of each field pair. Figure 4 The similarities shown are all semantic similarities corresponding to the main diagonal. Then, remove the fields with the highest semantic similarity in the same row, and then adjust the order of the remaining table fields or statement fields, such as Figure 4The semantic similarity of the two fields shown in the dotted box is significantly improved after adjustment, which means that there is an abnormal field order misalignment in the SQL statement. Figure 4 The two fields shown in the solid-line box have no improved semantic similarity after adjustment. If all other fields are removed or the semantic similarity does not improve after adjustment, the SQL statement will not be determined to have a field order mismatch. This avoids the problem of identifying equivalent fields due to different names as having a field order mismatch.
[0069] Based on the same concept, the embodiment of the present disclosure provides a data processing device based on a large model, such as Figure 5 As shown, the data processing device 500 includes: Determination module 501, for determining table fields in a target data table and statement fields in a target SQL statement, wherein the target SQL statement is used to write data to the target data table or update data in the target data table; A semantic understanding module 502 is configured to, when the number of the table fields and the number of the sentence fields are equal, perform semantic understanding on each pair of the table fields and the sentence fields using a large model to obtain a first semantic similarity matrix; an adjusting module 503, configured to adjust a field order of a first field according to the first semantic similarity matrix, and determine a second semantic similarity matrix corresponding to the adjusted first field and a second field, wherein the first field is the table field or the statement field, the second field is the table field or the statement field, and the first field and the second field are different; The recognition module 504 is configured to determine a recognition result of the target SQL statement based on the first semantic similarity matrix and the second semantic similarity matrix, wherein the recognition result is used to indicate whether the order of fields in the target SQL statement is normal or abnormal.
[0070] Optionally, the semantic similarity corresponding to the i-th row and j-th column in the first semantic similarity matrix represents the semantic similarity between the i-th field in the sentence field and the j-th field in the table field, where i and j are positive integers, and the adjustment module 503 is configured to: For each semantic similarity on the main diagonal of the first semantic similarity matrix, if the semantic similarity is the maximum value in the row, determining the semantic similarity as the target semantic similarity; Determining a third field corresponding to the target semantic similarity in the first field, and determining a first average value of remaining semantic similarities on the main diagonal except the target semantic similarity; When the first average value is less than a first preset threshold, the field order of the fourth field in the first field except the third field is adjusted.
[0071] Optionally, the adjustment module 503 is configured to: When the number of the fourth fields is less than or equal to a second preset threshold, traverse all possible field orders of the fourth fields to obtain at least one adjusted first field; respectively determining semantic similarity matrices corresponding to the adjusted first fields and the second fields to obtain a plurality of semantic similarity matrices; Among the multiple semantic similarity matrices, a semantic similarity matrix having the largest average semantic similarity corresponding to the fourth field on the main diagonal is determined as the second semantic similarity matrix.
[0072] Optionally, the adjustment module 503 is configured to: When the number of the fourth fields is greater than a second preset threshold, randomly adjusting the field order of the fourth fields in the first field to obtain a fifth field, until the average semantic similarity corresponding to the fourth fields on the main diagonal of the semantic similarity matrix corresponding to the fifth field and the second field is greater than the first average value; A semantic similarity matrix corresponding to the fifth field and the second field is determined as the second semantic similarity matrix.
[0073] Optionally, the identification module 504 is configured to: Determining a second average value of the semantic similarities corresponding to the fourth field on the main diagonal of the first semantic similarity matrix, and determining a third average value of the semantic similarities corresponding to the fourth field on the main diagonal of the second semantic similarity matrix; A recognition result of the target SQL statement is determined based on the second average value and the third average value.
[0074] Optionally, the identification module 504 is configured to: If the difference between the third average value and the second average value is greater than a third preset threshold, determining an identification result indicating that the target SQL statement has a field order misalignment anomaly; or When a ratio of a difference obtained by subtracting the second average value from the third average value to the second average value is greater than a fourth preset threshold, an identification result indicating that a field order misalignment abnormality exists in the target SQL statement is determined.
[0075] Optionally, the data processing device 500 further includes a parsing module, and the parsing module is configured to: Parsing the target SQL statement to obtain first information representing processing logic of fields in the target SQL statement, and parsing table structure information of the target data table to obtain second information, wherein the second information includes a table name of the target data table and / or field description information in the target data table; The semantic understanding module 502 is used to: The first semantic similarity matrix is obtained by performing semantic understanding on each pair of fields in the table field and the statement field based on the first information, the second information, the field name of the table field and the field name of the statement field through the big model.
[0076] Optionally, the determining module 501 is configured to: Parsing the table structure information of the target data table to obtain a first original field corresponding to the target data table, parsing the target SQL statement to obtain a second original field corresponding to the target SQL statement; Determine a target field having the same arrangement sequence number and the same field name in the first original field and the second original field; The table field is obtained by removing the target field from the first original field, and the statement field is obtained by removing the target field from the second original field.
[0077] Optionally, the data processing device 500 further includes a preprocessing module, and the preprocessing module is configured to: Perform at least one of the following preprocessing on the second original field: If the second original field has a third original field without a specified field name, completing the field name of the third original field based on the processing logic of the third original field in the target SQL statement; performing standardization processing on the field name of the second original field, wherein the standardization processing includes at least one of unifying uppercase and lowercase letters and removing special characters; Correct the field name with spelling errors in the second original field; The determining module 501 is used to: Determine target fields having the same arrangement sequence number and the same field name in the first original field and the preprocessed second original field.
[0078] Based on the same concept, an embodiment of the present disclosure further provides a computer-readable medium on which a computer program is stored. When the program is executed by a processing device, the program implements the steps of any of the above-mentioned large model-based data processing methods.
[0079] Based on the same concept, an embodiment of the present disclosure further provides an electronic device, which may include: a storage device having a computer program stored thereon; A processing device is used to execute the computer program in the storage device to implement the steps of any of the above-mentioned large model-based data processing methods.
[0080] Based on the same concept, an embodiment of the present disclosure also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned large model-based data processing methods when executed by a processor.
[0081] Reference below Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0082] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or programs loaded from a storage device 608 into a random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0083] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Figure 6 The electronic device 600 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0084] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0085] It should be noted that the computer-readable medium described above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0086] In some embodiments, communications may be conducted using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0087] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0088] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: determines a table field in a target data table and a statement field in a target SQL statement, wherein the target SQL statement is used to write data to the target data table or update data in the target data table; when the number of the table fields and the statement fields is equal, performs semantic understanding on each pair of the table fields and the statement fields using a large model to obtain a first semantic similarity matrix; adjusts the field order of the first field according to the first semantic similarity matrix, and determines a second semantic similarity matrix corresponding to the adjusted first field and the second field, wherein the first field is the table field or the statement field, and the second field is the table field or the statement field, and the first field and the second field are different; determines a recognition result of the target SQL statement according to the first semantic similarity matrix and the second semantic similarity matrix, wherein the recognition result is used to characterize whether the field order of the target SQL statement is normal or abnormally misplaced.
[0089] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0091] The modules described in the embodiments of the present disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the module itself.
[0092] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0093] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the present disclosure is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the scope of the above disclosure. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0095] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0096] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. A data processing method based on a large model, characterized in that: The data processing method includes: Determining a table field in a target data table and a statement field in a target SQL statement, wherein the target SQL statement is used to write data to the target data table or update data in the target data table; When the number of the table fields and the number of the sentence fields are equal, semantic understanding is performed on each pair of the table fields and the sentence fields using the large model to obtain a first semantic similarity matrix; Adjusting a field order of a first field according to the first semantic similarity matrix, and determining a second semantic similarity matrix corresponding to the adjusted first field and a second field, wherein the first field is the table field or the statement field, the second field is the table field or the statement field, and the first field and the second field are different; A recognition result of the target SQL statement is determined according to the first semantic similarity matrix and the second semantic similarity matrix, where the recognition result is used to characterize whether the order of fields in the target SQL statement is normal or abnormal.
2. The data processing method based on a large model according to claim 1, characterized in that: The semantic similarity corresponding to the i-th row and j-th column in the first semantic similarity matrix represents the semantic similarity between the i-th field in the sentence field and the j-th field in the table field, where i and j are positive integers. Adjusting the field order of the first field according to the first semantic similarity matrix includes: For each semantic similarity on the main diagonal of the first semantic similarity matrix, if the semantic similarity is the maximum value in the row, determining the semantic similarity as the target semantic similarity; Determining a third field corresponding to the target semantic similarity in the first field, and determining a first average value of remaining semantic similarities on the main diagonal except the target semantic similarity; When the first average value is less than a first preset threshold, the field order of the fourth field in the first field except the third field is adjusted.
3. The data processing method based on a large model according to claim 2, characterized in that: The adjusting the field order of the fourth field in the first field except the third field includes: When the number of the fourth fields is less than or equal to a second preset threshold, traverse all possible field orders of the fourth fields to obtain at least one adjusted first field; The determining of the second semantic similarity matrix corresponding to the adjusted first field and the second field includes: respectively determining semantic similarity matrices corresponding to the adjusted first fields and the second fields to obtain a plurality of semantic similarity matrices; Among the multiple semantic similarity matrices, a semantic similarity matrix having the largest average semantic similarity corresponding to the fourth field on the main diagonal is determined as the second semantic similarity matrix.
4. The data processing method based on a large model according to claim 2, characterized in that: The adjusting the field order of the fourth field in the first field except the third field includes: When the number of the fourth fields is greater than a second preset threshold, randomly adjusting the field order of the fourth fields in the first field to obtain a fifth field, until the average semantic similarity corresponding to the fourth fields on the main diagonal of the semantic similarity matrix corresponding to the fifth field and the second field is greater than the first average value; The determining of the second semantic similarity matrix corresponding to the adjusted first field and the second field includes: A semantic similarity matrix corresponding to the fifth field and the second field is determined as the second semantic similarity matrix.
5. The data processing method based on a large model according to claim 2, characterized in that: Determining a recognition result of the target SQL statement according to the first semantic similarity matrix and the second semantic similarity matrix includes: Determining a second average value of the semantic similarities corresponding to the fourth field on the main diagonal of the first semantic similarity matrix, and determining a third average value of the semantic similarities corresponding to the fourth field on the main diagonal of the second semantic similarity matrix; A recognition result of the target SQL statement is determined based on the second average value and the third average value.
6. The data processing method based on a large model according to claim 5, characterized in that: The determining, based on the second average value and the third average value, a recognition result of the target SQL statement includes: If the difference between the third average value and the second average value is greater than a third preset threshold, determining an identification result indicating that the target SQL statement has a field order misalignment anomaly; or When a ratio of a difference obtained by subtracting the second average value from the third average value to the second average value is greater than a fourth preset threshold, an identification result indicating that a field order misalignment abnormality exists in the target SQL statement is determined.
7. The data processing method based on a large model according to any one of claims 1 to 6, characterized in that: The data processing method further includes: Parsing the target SQL statement to obtain first information representing processing logic of fields in the target SQL statement, and parsing table structure information of the target data table to obtain second information, wherein the second information includes a table name of the target data table and / or field description information in the target data table; The method of performing semantic understanding on each pair of the table fields and the sentence fields using the large model to obtain a first semantic similarity matrix includes: The first semantic similarity matrix is obtained by performing semantic understanding on each pair of fields in the table field and the sentence field based on the first information, the second information, the field name of the table field and the field name of the sentence field.
8. The data processing method based on a large model according to any one of claims 1 to 6, characterized in that: Determining the table fields in the target data table and the statement fields in the target SQL statement includes: Parsing the table structure information of the target data table to obtain a first original field corresponding to the target data table, parsing the target SQL statement to obtain a second original field corresponding to the target SQL statement; Determine a target field having the same arrangement sequence number and the same field name in the first original field and the second original field; The table field is obtained by removing the target field from the first original field, and the statement field is obtained by removing the target field from the second original field.
9. The data processing method based on a large model according to claim 8, characterized in that: The data processing method further includes: Perform at least one of the following preprocessing on the second original field: If the second original field has a third original field without a specified field name, completing the field name of the third original field based on the processing logic of the third original field in the target SQL statement; performing standardization processing on the field name of the second original field, wherein the standardization processing includes at least one of unifying uppercase and lowercase letters and removing special characters; Correct the field name with spelling errors in the second original field; The determining of the target fields having the same arrangement sequence number and the same field name in the first original field and the second original field includes: Determine target fields having the same arrangement sequence number and the same field name in the first original field and the preprocessed second original field.
10. A data processing device based on a large model, characterized in that: The data processing device includes: a determination module, configured to determine table fields in a target data table and statement fields in a target SQL statement, wherein the target SQL statement is used to write data to the target data table or update data on the target data table; a semantic understanding module, configured to, when the number of the table fields and the number of the sentence fields are equal, perform semantic understanding on each pair of the table fields and the sentence fields using a large model to obtain a first semantic similarity matrix; an adjustment module, configured to adjust a field order of a first field according to the first semantic similarity matrix, and determine a second semantic similarity matrix corresponding to the adjusted first field and a second field, wherein the first field is the table field or the statement field, the second field is the table field or the statement field, and the first field and the second field are different; The recognition module is used to determine a recognition result of the target SQL statement based on the first semantic similarity matrix and the second semantic similarity matrix, wherein the recognition result is used to characterize whether the order of fields in the target SQL statement is normal or abnormal.
11. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 9 are implemented.
12. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Method and device for consistency testing of field sequence
CN106649333A
Text similarity determination method and device
CN115827818A
Text similarity calculation method and device and electronic equipment
CN117688399A
SQL statement detection method and device, equipment and medium
CN119474125A
Method and system for determining similarity of items based on similarity objects and their features
US20060112068A1