Vehicle model data matching method and device, and electronic equipment
Patent Information
- Application Number
- CN202211101754.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-09
AI Technical Summary
[0005]有鉴于此,本申请提供了一种车型数据的匹配方法、装置及电子设备,主要目的在于改善目前现有的硬匹配方法,对于实体名称字段顺序不同但实质内容相同的车型数据,无法做到车型数据准确匹配的技术问题
[0018]借由上述技术方案,本申请提供的一种车型数据的匹配方法、装置及电子设备,与目前现有的硬匹配方法相比,本申请对于实体名称字段顺序不同但实质内容相同的车型数据,也能做到车型数据的准确匹配,提升了匹配的召回率与精准度,具体可首先将待匹配的车型数据以汽车特征为子粒度进行拆分,得到车型数据的多个基本特征元素;再将多个第一基本特征元素按照不同顺序组合后得到的车型数据,分别与目标数据库中的各个车型数据进行匹配;然后获取目标数据库中匹配率大于或等于目标匹配率阈值的车型数据,作为待匹配的车型数据的匹配结果。通过应用本申请的技术方案,将车型数据拆分并进行无序的子粒度匹配,使得虽然车型数据中各实体名称字段的排列顺序不同,但是也能精准匹配到含义实质相同的两条车型数据,进而有效提高数据之间的车型匹配精准度,使得尽可能地匹配到目标数据库中的车型数据,以利用其进行汽车相关数据分析与预测,从而可从数据智能运营角度支撑车辆运营精准决策。
Smart Images

Figure CN117688392B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a method, apparatus, and electronic device for matching vehicle model data. Background Technology
[0002] With the development of automotive big data, the demand for accurate vehicle model databases is constantly increasing. It is necessary not only to process simple data, but also to control the quality of vehicle model data under the premise that the amount of data is increasing exponentially.
[0003] Currently, a hard matching method can be used for vehicle model matching. Specifically, the vehicle model data to be matched is matched with the vehicle model data in the database, and the matching data from the database is the same as the vehicle model data to be matched.
[0004] However, this hard matching method can only be applied to the matching process of completely identical vehicle model data. It cannot achieve accurate matching of vehicle model data for vehicle model data with different entity name field orders but the same substantive content. Summary of the Invention
[0005] In view of this, this application provides a method, apparatus and electronic device for matching vehicle model data. The main purpose is to improve the existing hard matching method, which cannot accurately match vehicle model data with different entity name field orders but the same substantive content.
[0006] Firstly, this application provides a method for matching vehicle model data, including:
[0007] Obtain vehicle model data to be matched;
[0008] The vehicle model data to be matched is split into multiple first basic feature elements by using vehicle features as the sub-granularity.
[0009] The vehicle model data obtained by combining multiple first basic feature elements in different orders is matched with each vehicle model data in the target database.
[0010] The vehicle model data with a matching rate greater than or equal to the target matching rate threshold in the target database is obtained as the matching result of the vehicle model data to be matched.
[0011] Secondly, this application provides a vehicle model data matching device, comprising:
[0012] The acquisition module is configured to acquire vehicle model data to be matched;
[0013] The splitting module is configured to split the vehicle model data to be matched into multiple first basic feature elements of the vehicle model data by dividing it into vehicle features as sub-granularities.
[0014] The matching module is configured to match the vehicle model data obtained by combining multiple first basic feature elements in different orders with the vehicle model data in the target database respectively.
[0015] The acquisition module is configured to acquire vehicle model data in the target database whose matching rate is greater than or equal to the target matching rate threshold, and use this data as the matching result of the vehicle model data to be matched.
[0016] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the vehicle model data matching method described in the first aspect.
[0017] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the vehicle model data matching method described in the first aspect.
[0018] By employing the above technical solution, this application provides a vehicle model data matching method, apparatus, and electronic device. Compared with existing hard matching methods, this application can accurately match vehicle model data even if the order of entity name fields is different but the substantive content is the same, thus improving the recall and precision of the matching. Specifically, the vehicle model data to be matched is first split into multiple basic feature elements based on vehicle features. Then, the vehicle model data obtained by combining multiple basic feature elements in different orders is matched with each vehicle model data in the target database. Finally, the vehicle model data with a matching rate greater than or equal to the target matching rate threshold in the target database is obtained as the matching result of the vehicle model data to be matched. By applying the technical solution of this application, the vehicle model data is split and matched at unordered sub-granularity, so that even if the order of the entity name fields in the vehicle model data is different, two vehicle model data with the same substantive meaning can still be accurately matched. This effectively improves the accuracy of vehicle model matching between data, enabling matching with vehicle model data in the target database as much as possible. This allows for the use of vehicle-related data analysis and prediction, thereby supporting accurate decision-making in vehicle operation from the perspective of data-driven intelligent operation.
[0019] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a vehicle model data matching method provided in an embodiment of this application is shown.
[0023] Figure 2 A schematic diagram illustrating an example of vehicle model data splitting provided in an embodiment of this application is shown;
[0024] Figure 3 A flowchart illustrating another vehicle model data matching method provided in an embodiment of this application is shown;
[0025] Figure 4 This illustration shows an example diagram of synonym expansion provided in an embodiment of this application;
[0026] Figure 5 This paper presents an example schematic diagram of a vehicle model data matching scenario provided in an embodiment of this application.
[0027] Figure 6 A schematic diagram of a vehicle model data matching device provided in an embodiment of this application is shown. Detailed Implementation
[0028] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0029] To address the technical problem of existing hard-matching methods failing to accurately match vehicle model data even when the entity name fields have different orders but the actual content is the same, this embodiment provides a vehicle model data matching method, such as... Figure 1 As shown, the method includes:
[0030] Step 101: Obtain vehicle model data to be matched.
[0031] In this embodiment, it is necessary to match the vehicle model data to be matched with the vehicle model data in the target database, and then find the matching vehicle model data from the target database as the matching result of the vehicle model data to be matched.
[0032] For example, given the vehicle model data to be matched: 【2019 A Brand B1 Series 1.6MT Standard 7-seater National VI】, we need to find the matching model from the massive amount of vehicle data in the target database. The vehicle model data in the target database can be seen as follows:
[0033] [2019 B1 Series 2.0L 9-seater Comfort Model, 2019 B1 Series 2.0L 7-seater Comfort Model, 2019 B2 Series 1.6L Special Edition 6-seater, 2019 B2 Series 1.6L Special Edition 7-seater, 2019 B3 Series 2.0L 9-seater Standard Model, ..., 2019 Facelifted B1 Series 1.6L 7-seater Basic Model (China VI Emission Standard), 2019 B1 Series 1.6L 7-seater Comfort Model (China VI Emission Standard), 2019 C Series 1.6L 2-seater Standard Model (China VI Emission Standard)...]
[0034] Find the model data that uniquely corresponds to the model data to be matched in the target database: 【2019 B1 Series 1.6L 7-seater Standard Model National VI】.
[0035] To achieve the above objective, the process shown in steps 102 to 104 can be performed.
[0036] Step 102: Split the vehicle model data to be matched into sub-granularities based on vehicle features to obtain multiple first basic feature elements of the vehicle model data to be matched.
[0037] This embodiment constructs the basic feature elements of vehicle model data, decomposing unstructured vehicle model data into structured feature data with vehicle features as sub-granularities. Specifically, natural language processing (NLP) techniques can be used for this decomposition. NLP is an important field in computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language.
[0038] For example, such as Figure 2As shown, the vehicle model data to be matched can be broken down into sub-granularities based on vehicle characteristics such as "drive type," "model year," "brand," "version," "model series," "energy," "environmental standard," "seat count," "system," "transmission," and "displacement," resulting in the basic feature elements (first basic feature elements) of the vehicle model data to be matched. Specifically, "drive type" can include: two-wheel drive, four-wheel drive, etc.; "model year" can include: 2021 model, 2022 model, etc.; "version" can include: premium version, sports version, etc.; "energy" can include: range-extended, gasoline, etc.; "environmental standard" (regulations on vehicle emissions) can include: China V, China VI, etc.; "seat count" can include: 5 seats, 7 seats, etc.; and "displacement" can include: 1.2T, 1.5T, etc.
[0039] Step 103: The vehicle model data obtained by combining multiple first basic feature elements in different orders is matched with the vehicle model data in the target database.
[0040] Unordered sub-granularity matching involves matching vehicle model data without considering the order of basic feature elements. The matching rate for each vehicle model data entry in the target database is determined by assessing whether it contains all the basic feature elements (obtained through splitting in step 102) of the vehicle model data to be matched. For example, the vehicle model data to be matched can be split into sub-granularities based on vehicle features such as "drive type," "model year," "brand," "version," "series," "energy," "environmental standards," "seat count," "system," "transmission," and "displacement," resulting in multiple basic feature elements. These basic feature elements are then combined in different orders to obtain vehicle model data, which is then matched against various vehicle model data entries in the target database.
[0041] Step 104: Obtain vehicle model data from the target database whose matching rate is greater than or equal to the target matching rate threshold, and use this data as the matching result for the vehicle model data to be matched.
[0042] The target matching rate threshold can be pre-set according to actual needs; for example, it can be 100% or 99%. For instance, if it's necessary to match vehicle model data (the vehicle model data to be matched) in database A with vehicle model data in database B (the target database), since database B may contain insurance information for specific vehicle models, matching the vehicle model data can associate the vehicle model data in database A with the corresponding insurance information. Specifically, the vehicle model data to be matched in database A can first be split into basic feature elements based on vehicle characteristics; then, based on these basic feature elements, unordered sub-granularity matching is performed between the vehicle model data to be matched and each vehicle model data in database B; finally, the vehicle model data in database B with a matching rate of 100% is obtained as the matching result for the vehicle model data to be matched.
[0043] Compared to existing hard matching methods, this embodiment can accurately match vehicle model data even when the order of entity name fields differs but the substantive content is the same, improving the recall and precision of the matching. By applying the technical solution of this embodiment, vehicle model data is split and subjected to unordered sub-granularity matching. This allows for accurate matching of two vehicle model data with essentially the same meaning, even if the order of the entity name fields in the vehicle model data is different. This effectively improves the accuracy of vehicle model matching between data, thereby supporting precise decision-making in vehicle operation from the perspective of data-driven intelligent operation.
[0044] Furthermore, as a refinement and extension of the above embodiments, in order to fully illustrate the specific implementation process of the method in this embodiment, this embodiment provides the following: Figure 3 The specific method shown includes:
[0045] Step 201: Obtain vehicle model data to be matched.
[0046] Step 202: Split the vehicle model data to be matched into sub-granularities based on vehicle features to obtain multiple first basic feature elements of the vehicle model data.
[0047] Step 203: The vehicle model data obtained by combining multiple first basic feature elements in different orders is matched with the vehicle model data in the target database.
[0048] In practice, the basic feature elements obtained from the decomposition may contain synonyms with completely identical meanings. Therefore, to achieve more comprehensive and accurate data matching, these basic feature elements can optionally be expanded with synonyms. Accordingly, step 203 may specifically include: firstly, expanding the first basic feature elements with synonyms; then, matching the vehicle model data obtained by combining the multiple first basic feature elements with synonyms in different orders with the vehicle model data in the target database.
[0049] For example, a thesaurus can be used to expand the basic feature elements obtained from these splits with synonyms, such as... Figure 4 As shown, "AT Deluxe Edition" can be divided into "AT" and "Deluxe Edition," and "Deluxe Edition" can be further divided into "Deluxe" and "Version." Through a thesaurus search, it is determined that "AT" and "Automatic" are synonyms, while "Version" and "Type" are synonyms. Therefore, the synonyms for "AT Deluxe Edition" are "Automatic Deluxe Type," "Automatic Deluxe Edition," etc. Based on these synonyms, the basic feature elements obtained from the decomposition are expanded with synonyms.
[0050] For example, all the basic feature elements (expanded with synonyms) in vehicle model data c (vehicle model data to be matched) are combined in different orders and compared with vehicle model data d (a vehicle model data in the target database) in the target database. The matching rate corresponding to the order combination with the highest matching rate is obtained as the corresponding matching rate.
[0051] The above methods can accurately match vehicle model data at disordered sub-granularity, thereby improving the accuracy of vehicle data matching.
[0052] Step 204a: If there are vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold, then obtain the vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold, and use it as the matching result of the vehicle model data to be matched.
[0053] In step 204b, which is parallel to step 204a, if there is no vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold, then multiple first basic feature elements and multiple second basic feature elements of each vehicle model data in the target database are calculated using a deep learning model to determine the semantic similarity between the vehicle model data to be matched and each vehicle model data.
[0054] The second basic feature element can be obtained by splitting the vehicle model data in the target database according to vehicle features as sub-granularity. The deep learning model can specifically be a BiLSTM+Attention model, which can be used to calculate the semantic similarity between vehicle model data. The deep learning model can be trained based on positive and negative samples. Optionally, the training process of the deep learning model can include: first, obtaining positive and negative sample data. Positive sample data can be vehicle model data that has been split according to vehicle features and then matched at an unordered sub-granularity in the target database. The matching rate obtained is greater than or equal to the target matching rate threshold, such as a direct matching result with a matching rate of 100%, which is used as the model's positive sample. Negative sample data can be vehicle model data randomly obtained from vehicle model data with a matching rate less than the target matching rate threshold, such as data other than the correctly matched results, randomly obtained as negative samples. Then, the model is trained based on the positive and negative sample data to obtain the deep learning model. During model training, manual annotation results can be combined for model verification, thereby training a qualified BiLSTM+Attention model.
[0055] In practical applications of the model, since semantic similarity calculation is resource-intensive, step 204b can optionally include: firstly, selecting vehicle model data that meets the recall feature conditions from the various vehicle model data; then, using a deep learning model, calculating the semantic similarity between the vehicle model data to be matched and the vehicle model data that meets the recall feature conditions using multiple first basic feature elements and multiple second basic feature elements of the vehicle model data that meets the recall feature conditions, to determine the semantic similarity between the vehicle model data to be matched and the vehicle model data that meets the recall feature conditions. This optional approach reduces the amount of data requiring semantic similarity calculation, effectively saving system resources and improving the efficiency of vehicle model data matching.
[0056] For example, the above-mentioned process of filtering vehicle data that meets the recall feature conditions from various vehicle data may specifically include: firstly, obtaining the vehicle features corresponding to multiple first basic feature elements of the vehicle data to be matched; then, determining the vehicle features as the third basic feature elements of the target vehicle features from the multiple first basic feature elements of the vehicle data to be matched; and then, filtering out vehicle data containing all the third basic feature elements from various vehicle data in the target database as vehicle data that meets the recall feature conditions.
[0057] The target vehicle features can be features such as "model year," "brand," and "version," which can be preset according to the actual situation. For example, relatively important vehicle features in vehicle data matching can be selected as target vehicle features, such as "model year," "brand," and "version," which better reflect the obvious characteristics of the vehicle data. During the vehicle model data matching process, these vehicle features can be used to filter out a more refined range of vehicle model data. From the multiple basic feature elements of the vehicle model data to be matched, the basic feature elements for "model year," "brand," and "version" are determined. These three basic feature elements are the third basic feature elements for the target vehicle feature.
[0058] For example, such as Figure 5 As shown, database A stores vehicle model data to be matched, and the corresponding vehicle model data needs to be matched in database B. First, all vehicle model data in both databases A and B can be split into sub-granularities based on vehicle features to obtain the basic feature elements of each vehicle model. Then, a single vehicle model data 'a' (containing the multiple basic feature elements obtained from the splitting) is retrieved from database A and further matched with vehicle model data in database B that meets the criteria of "model year," "brand," and "version" (selecting model year, brand, and version as strong recall features, i.e., filtering out vehicle model data that meets the recall feature conditions) using a deep learning model (i.e., calculating semantic similarity). For example, if vehicle model data 'a' has a model year of "2021," a brand of "XX," and a version of "Luxury Supreme Edition," then vehicle model data 'a' is matched with all vehicle model data in database B that contains "2021," "XX," and "Luxury Supreme Edition" based on semantic similarity.
[0059] This optional approach reduces the amount of data required for semantic similarity calculations, effectively saves system resources, and improves the efficiency of vehicle model data matching.
[0060] Step 205b: Obtain vehicle model data with semantic similarity greater than or equal to the target similarity threshold from each vehicle model data in the target database, and use this data as the matching result for the vehicle model data to be matched.
[0061] Based on the optional method in step 204b, step 205b may specifically include: obtaining vehicle model data with semantic similarity greater than or equal to a preset similarity threshold (such as 95% or 98%) from vehicle model data that meet the preset recall feature conditions, and using this as the matching result of the vehicle model data to be matched.
[0062] For example, based on such Figure 5The example shown retrieves a vehicle model data 'a' (containing multiple basic feature elements obtained from the breakdown) from database A and further matches it with vehicle model data from database B that meets the criteria of "model year", "brand", and "version" (selecting model year, brand, and version as strong recall features, i.e., filtering out vehicle model data that meets the preset recall feature conditions). The matched vehicle model data is then sorted in descending order of semantic similarity, and the vehicle model data with a semantic similarity greater than or equal to 98% is selected as the final matching result corresponding to vehicle model data 'a'.
[0063] Optionally, if the target database (e.g., a database of domestic vehicle model data) does not contain vehicle model data with a semantic similarity greater than or equal to the target similarity threshold, it indicates that the vehicle model data to be matched is likely foreign vehicle model data. In this embodiment, the method may further include: querying converted vehicle model data corresponding to the vehicle model data to be matched, wherein the vehicle corresponding to the converted vehicle model data is produced in the target production area (e.g., a domestic production area), and the target production area can be the production area of the vehicle (domestic vehicle) corresponding to the vehicle model data in the target database; then splitting the converted vehicle model data into multiple fourth basic feature elements of the vehicle model data using vehicle features as sub-granularity; then matching the vehicle model data obtained by combining the multiple fourth basic feature elements in different orders with each vehicle model data in the target database; and finally obtaining the vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold as the matching result of the vehicle model data to be matched.
[0064] For example, in practical applications, for some foreign-made vehicle model data, it is not possible to directly determine the matching vehicle model data in a domestic database through unordered sub-granularity matching and semantic similarity calculation. It is necessary to first convert the vehicle model data into converted vehicle model data of substantially the same domestic model, and then use the converted vehicle model data to perform unordered sub-granularity matching, so as to achieve accurate matching for this part of the vehicle model data.
[0065] By applying the technical solution of this embodiment, a vehicle model matching method combining vehicle model data segmentation and semantic parsing is provided to address the vehicle model data matching problem. This solves the recall and accuracy issues in vehicle model data matching, thereby supporting precise decision-making in vehicle operation from a data-driven intelligent operation perspective. Compared to manual annotation matching, this solution greatly improves efficiency and solves the problem of not being able to annotate massive amounts of one-to-many text. Compared to hard matching, this solution can also achieve accurate matching of vehicle model data even when the order of entity name fields is different but the substantive content is the same. Compared to synonym replacement matching, this solution solves the problems of incomplete data and differences in contextual word order, effectively addressing the low recall rate issue.
[0066] Furthermore, as Figure 1 and Figure 3 The specific implementation of the method shown in this embodiment provides a vehicle model data matching device, such as... Figure 6 As shown, the device includes: an acquisition module 31, a splitting module 32, and a matching module 33.
[0067] The acquisition module 31 is configured to acquire vehicle model data to be matched;
[0068] The splitting module 32 is configured to split the vehicle model data to be matched into multiple first basic feature elements of the vehicle model data by using vehicle features as the sub-granularity.
[0069] The matching module 33 is configured to match the vehicle model data obtained by combining multiple first basic feature elements in different orders with the vehicle model data in the target database respectively.
[0070] The acquisition module 31 is configured to acquire vehicle model data in the target database whose matching rate is greater than or equal to the target matching rate threshold, and use this data as the matching result of the vehicle model data to be matched.
[0071] In a specific application scenario, the matching module 33 is specifically configured to expand the first basic feature element with synonyms; and to match the vehicle model data obtained by combining the multiple first basic feature elements after synonym expansion in different orders with the vehicle model data in the target database.
[0072] In specific application scenarios, the matching module 33 is further configured to, if there is no vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold, calculate the semantic similarity between the vehicle model data to be matched and each of the vehicle model data in the target database using a deep learning model, wherein the second basic feature element is obtained by splitting each vehicle model data with vehicle features as the sub-granularity; and obtain vehicle model data with a semantic similarity greater than or equal to the target similarity threshold from each of the vehicle model data as the matching result of the vehicle model data to be matched.
[0073] In specific application scenarios, optionally, the preset deep learning model is trained based on positive and negative sample data. The positive sample data consists of vehicle model data that has been split according to vehicle features and then matched at an unordered sub-granularity level in the target database, with the resulting matching rate being greater than or equal to the target matching rate threshold. The negative sample data consists of vehicle model data randomly obtained from the vehicle model data with a matching rate less than the target matching rate threshold.
[0074] In a specific application scenario, the matching module 33 is further configured to: filter out vehicle model data that meets the recall feature conditions from the various vehicle model data; calculate the semantic similarity between the vehicle model data to be matched and the vehicle model data that meets the recall feature conditions using a deep learning model; and obtain vehicle model data with a semantic similarity greater than or equal to a target similarity threshold from the vehicle model data that meets the recall feature conditions, as the matching result of the vehicle model data to be matched.
[0075] In a specific application scenario, the matching module 33 is further configured to obtain the vehicle features corresponding to the multiple first basic feature elements of the vehicle model data to be matched; determine the vehicle feature as the third basic feature element of the target vehicle feature from the multiple first basic feature elements; and filter out the vehicle model data containing all the third basic feature elements from the various vehicle model data in the target database as the vehicle model data that meets the recall feature conditions.
[0076] In a specific application scenario, the matching module 33 is further configured to: after calculating the semantic similarity between the vehicle model data to be matched and each of the vehicle model data in the target database using a deep learning model, if there is no vehicle model data in the target database with a semantic similarity greater than or equal to a target similarity threshold, query the converted vehicle model data corresponding to the vehicle model data to be matched, wherein the vehicle corresponding to the converted vehicle model data is produced in a target production area, and the target production area is the production area of the vehicle corresponding to the vehicle model data in the target database; split the converted vehicle model data into multiple fourth basic feature elements of the vehicle model data using vehicle features as sub-granularity; match the vehicle model data obtained by combining multiple fourth basic feature elements in different orders with each of the vehicle model data in the target database; and obtain the vehicle model data in the target database with a matching rate greater than or equal to a target matching rate threshold as the matching result of the vehicle model data to be matched.
[0077] It should be noted that other corresponding descriptions of the functional units involved in the vehicle model data matching device provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 3 The corresponding description in [the document] will not be repeated here.
[0078] Based on the above, Figure 1 and Figure 3Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 3 The method shown.
[0079] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0080] Based on the above, Figure 1 and Figure 3 The method shown, and Figure 6 To achieve the above objectives, this application also provides an electronic device, specifically a personal computer, server, etc., as shown in the virtual device embodiment. The device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 3 The method shown.
[0081] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0082] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0083] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented using software plus necessary general-purpose hardware platforms, or it can be implemented using hardware. The solution of this embodiment can be applied to the field of automotive information mining and matching technology, specifically a short text matching method in the vertical automotive field, and can be applied to automotive-related data analysis and prediction. Compared with existing technologies, it can accurately match vehicle model data even for vehicle model data with different entity name field orders but the same substantive content, improving the recall rate and accuracy of the matching. By applying the technical solution of this embodiment, vehicle model data is split and unordered sub-granularity matching is performed, so that even if the order of the entity name fields in the vehicle model data is different, two vehicle model data with the same substantive meaning can still be accurately matched, thereby effectively improving the accuracy of vehicle model matching between data, and thus supporting accurate vehicle operation decision-making from the perspective of data-driven intelligent operation.
[0085] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0086] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for matching vehicle model data, characterized in that, include: Obtain vehicle model data to be matched; The vehicle model data to be matched is split into multiple first basic feature elements by using vehicle features as the sub-granularity. The vehicle model data obtained by combining multiple first basic feature elements in different orders is matched with each vehicle model data in the target database. Obtain vehicle model data from the target database whose matching rate is greater than or equal to the target matching rate threshold, and use this data as the matching result for the vehicle model data to be matched. After the vehicle model data obtained by combining multiple first basic feature elements in different orders is matched with each vehicle model data in the target database, the method further includes: If there is no vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold, then multiple first basic feature elements and multiple second basic feature elements of each vehicle model data in the target database are calculated using a deep learning model to determine the semantic similarity between the vehicle model data to be matched and each vehicle model data, wherein the second basic feature elements are obtained by splitting each vehicle model data with vehicle features as the sub-granularity. Vehicle data with semantic similarity greater than or equal to the target similarity threshold are obtained from the various vehicle data and used as the matching result of the vehicle data to be matched.
2. The method according to claim 1, characterized in that, The process of combining multiple first basic feature elements in different orders to obtain vehicle model data, and then matching them with vehicle model data in the target database, includes: Expand the first basic feature element with synonyms; The vehicle model data obtained by combining multiple first basic feature elements expanded with synonyms in different orders is matched with each vehicle model data in the target database.
3. The method according to claim 1, characterized in that, The training process of the deep learning model includes: Positive sample data and negative sample data are obtained, wherein the positive sample data are vehicle model data that are split according to vehicle features and then matched at an unordered sub-granularity level in the target database, and the resulting vehicle model data has a matching rate greater than or equal to a target matching rate threshold; the negative sample data are vehicle model data randomly obtained from the vehicle model data with a matching rate less than the target matching rate threshold. The deep learning model is obtained by training the model based on the positive sample data and the negative sample data.
4. The method according to claim 3, characterized in that, The step of calculating the semantic similarity between the vehicle model data to be matched and each of the vehicle model data in the target database using a deep learning model includes: From the data of each vehicle model, select the vehicle model data that meets the recall criteria; Multiple first basic feature elements and multiple second basic feature elements of the vehicle model data that meet the recall feature conditions are calculated using a deep learning model to determine the semantic similarity between the vehicle model data to be matched and the vehicle model data that meet the recall feature conditions. The step of obtaining vehicle model data with semantic similarity greater than or equal to the target similarity threshold from the various vehicle model data, and using this as the matching result for the vehicle model data to be matched, includes: From the vehicle model data that meets the recall feature conditions, obtain vehicle model data with semantic similarity greater than or equal to the target similarity threshold, and use this as the matching result of the vehicle model data to be matched.
5. The method according to claim 4, characterized in that, The step of filtering vehicle data that meets the recall criteria from the various vehicle data includes: Obtain the vehicle features corresponding to multiple first basic feature elements of the vehicle model data to be matched; From the plurality of first basic feature elements, determine the vehicle feature as the third basic feature element of the target vehicle feature; From the vehicle data of each vehicle model in the target database, vehicle data containing all the third basic feature elements are selected as vehicle data that meet the recall feature conditions.
6. The method according to claim 1, characterized in that, After calculating the semantic similarity between the vehicle model data to be matched and each of the vehicle model data in the target database using a deep learning model, the method further includes: If there is no vehicle model data in the target database with a semantic similarity greater than or equal to the target similarity threshold, then the converted vehicle model data corresponding to the vehicle model data to be matched is queried, wherein the vehicle corresponding to the converted vehicle model data is produced in the target production area, and the target production area is the production area of the vehicle corresponding to the vehicle model data in the target database; The converted vehicle model data is split into multiple fourth basic feature elements by using vehicle features as the sub-granularity. The vehicle model data obtained by combining multiple fourth basic feature elements in different orders is matched with each vehicle model data in the target database. The vehicle model data with a matching rate greater than or equal to the target matching rate threshold in the target database is obtained as the matching result of the vehicle model data to be matched.
7. A vehicle model data matching device, characterized in that, include: The acquisition module is configured to acquire vehicle model data to be matched; The splitting module is configured to split the vehicle model data to be matched into multiple first basic feature elements of the vehicle model data by dividing it into vehicle features as sub-granularities. The matching module is configured to match the vehicle model data obtained by combining multiple first basic feature elements in different orders with the vehicle model data in the target database respectively. The acquisition module is configured to acquire vehicle model data in the target database whose matching rate is greater than or equal to the target matching rate threshold, and use this data as the matching result of the vehicle model data to be matched. After the vehicle model data obtained by combining multiple first basic feature elements in different orders is matched with each vehicle model data in the target database, the device further includes: If there is no vehicle model data in the target database with a matching rate greater than or equal to the target matching rate threshold, then multiple first basic feature elements and multiple second basic feature elements of each vehicle model data in the target database are calculated using a deep learning model to determine the semantic similarity between the vehicle model data to be matched and each vehicle model data, wherein the second basic feature elements are obtained by splitting each vehicle model data with vehicle features as the sub-granularity. Vehicle data with semantic similarity greater than or equal to the target similarity threshold are obtained from the various vehicle data and used as the matching result of the vehicle data to be matched.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Precise automobile model matching system based on artificial intelligence algorithm
CN110555024A
Tablet identification method, readable storage medium and electronic equipment
CN114168772A