Treatment method
By obtaining and evaluating the target parameters of the delivery address editing operation and using multi-dimensional features to dynamically identify abnormal address changes, the problem of gray industries bypassing rules is solved, and the security and stability of the e-commerce platform is improved.
Patent Information
- Application Number
- CN202510900857.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
In the existing technology, gray industries make improper profits by changing the delivery address to circumvent the merchant interception rules, resulting in reduced security and stability of e-commerce platforms. Static rules are difficult to cope with the ever-changing operating modes of gray industries.
By obtaining the target parameters of the delivery address editing operation, including editing information and operation parameters, the credibility of the operation is dynamically evaluated. By using multi-dimensional features such as semantic difference, spatial similarity and operation frequency, it is determined whether the operation is abnormal and abnormal address changes are identified in real time.
It improves the accuracy of identifying abnormal address changes, reduces the possibility of shipping to risky addresses, and enhances the security and stability of the e-commerce platform.
Smart Images

Figure CN120805922A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to, but is not limited to, the field of information technology, and in particular to a processing method. BACKGROUND
[0002] In the field of electronic commerce, the delivery address can be used as important information for identifying the identity of a user.
[0003] In related technologies, there is a problem of gray production bypassing the interception rules of a merchant on a delivery address by changing the delivery address to make improper profits, which reduces the security and stability of an e-commerce platform. SUMMARY
[0004] Therefore, at least one processing method is provided in the embodiments of the present application.
[0005] The technical solutions of the embodiments of the present application are implemented as follows:
[0006] The embodiments of the present application provide a processing method, which includes:
[0007] In response to receiving a first operation, a target parameter corresponding to the first operation is obtained; the first operation includes an editing operation on a first delivery address, and the target parameter includes editing information of the first delivery address and / or an operation parameter of the first operation;
[0008] Based on the target parameter, a target credibility of the first operation is determined; the target credibility is used to determine whether the first operation is abnormal.
[0009] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the technical solutions of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 An implementation flowchart of a processing method provided by the embodiments of the present application Figure 1 ;
[0011] Figure 2 An implementation flowchart of a processing method provided by the embodiments of the present application Figure 2 ;
[0012] Figure 3 A component structure diagram of a processing device provided by the embodiments of the present application;
[0013] Figure 4 A component structure diagram of an electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION
[0014] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further described in detail below in combination with the drawings and embodiments, and the described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0015] In the following description, the terms "first / second / third" are merely to distinguish similar objects and do not represent a specific order of the objects. Understandably, the "first / second / third" can be interchanged in a specific order or sequence as permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs. The terms used herein are merely for the purpose of describing the present application and are not intended to limit the present application. In addition, it should be noted that only parts related to the present application are shown in the drawings for ease of description.
[0016] In the prior art, the delivery address is usually monitored based on static rules, such as limiting the order placing times of the same or similar address, intercepting high-risk addresses, etc. However, such methods mainly rely on static rules and are difficult to cope with the constantly changing operation modes of gray production, and lack the ability to dynamically evaluate the address editing operation itself. Therefore, a method capable of dynamically identifying abnormal address changes is needed to solve the above problems.
[0017] The embodiment of the present application provides a processing method, which can be executed by an electronic device. The electronic device can be a server, a notebook computer, a tablet computer, a desktop computer, a smart television, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated message device, a portable game device) and the like with data processing capability. As shown in the figure, the processing method comprises the following steps S101 to S102: Figure 1
[0018] Step S101, in response to receiving a first operation, obtaining a target parameter corresponding to the first operation; the first operation comprises an editing operation on a first delivery address, and the target parameter comprises editing information of the first delivery address and / or operation parameters of the first operation.
[0019] The delivery address is address information of a user receiving goods.
[0020] Exemplarily, the delivery address can include administrative region information corresponding to administrative region levels such as province, city, district, street, and coding information such as building number and house number.
[0021] The first operation is an operation of changing the delivery address of the user, and can include at least one of an editing operation such as adding, deleting, and modifying.
[0022] For example, the first operation can be an operation of adding a delivery address by the user.
[0023] For another example, the first operation can be an operation of modifying a character in an existing delivery address by the user.
[0024] For another example, the first operation can be an operation of deleting an existing delivery address by the user.
[0025] In some embodiments, the first operation can be a single editing operation of editing the first delivery address until the second delivery address is obtained. Wherein, each time the delivery address is updated, it can be regarded as an editing operation.
[0026] In some embodiments, the first operation can also be a sequence of editing actions of performing editing actions on characters in the first delivery address. Wherein, the editing actions can include but are not limited to adding actions, deleting actions, or modifying actions.
[0027] In some embodiments, the delivery address can be parsed and standardized before obtaining the target parameters corresponding to the first operation in response to receiving the first operation, that is, the original address input by the user is converted into structured data that can be analyzed. In this way, the consistency and reliability of the data can be improved, and the processing efficiency and accuracy of the subsequent processing can be further improved.
[0028] In some embodiments, the original first delivery address corresponding to the first operation and the second delivery address obtained after the editing operation on the first delivery address can be subjected to field structuring processing after receiving the first operation.
[0029] For example, illegal characters in the delivery address can be removed first, and then word segmentation processing and entity recognition can be performed.
[0030] For example, a pre-trained language model (Bidirectional Encoder Representations from Transformers, BERT model) can be used to identify address elements, and after distinguishing fields such as province, city, district, street, building, floor, and house number, the corresponding values for each field can be further matched.
[0031] In some embodiments, the delivery address obtained after field structuring processing can be subjected to standardization processing. In this way, the consistency of the address data can be further improved based on the processing standard.
[0032] For example, the specific information corresponding to the administrative region level in the address can be completed. For example, the information corresponding to the highest administrative region level can be automatically completed through the upper and lower association between the administrative region levels. For "Chaoyang District Wangjing Street", since "Chaoyang District" containing "Wangjing Street" is subordinate to "Beijing City", the address can be completed as "Beijing City Chaoyang District Wangjing Street".
[0033] For another example, the Point of Interest (POI) level of the delivery address can be normalized to establish a unified expression of the address dictionary. For example, the words such as "building", "unit", "floor", and "room" can be uniformly expressed in all addresses, thereby improving the efficiency and accuracy of substantive detection of the address.
[0034] In some embodiments, the geographic code corresponding to each delivery address, i.e., the code data representing the geographic spatial information of each delivery address, can be obtained. The geographic code can include, but is not limited to, at least one of latitude and longitude coordinates, planar rectangular coordinates, height, etc.
[0035] For example, the latitude and longitude coordinates corresponding to the delivery address can be obtained by calling the application interface of the map vendor.
[0036] The target parameter corresponding to the first operation can represent the operation behavior information of the first operation itself, the information related to the change of the first delivery address during the execution of the first operation, and / or the information related to the second delivery address obtained after the execution of the first operation. Therefore, by obtaining the target parameter corresponding to the first operation, the behavior characteristics of the editing operation performed by the user on the first delivery address can be better captured, thereby providing a data basis for subsequent operation evaluation and further improving the accuracy of identifying abnormal address change operations. The editing information of the first delivery address can include editing process information (i.e., information related to the change of the first delivery address during the execution of the first operation) and / or editing result information (i.e., information related to the second delivery address obtained after the execution of the first operation).
[0037] The operation parameter of the first operation can be used to represent the operation behavior information of the first operation itself. For example, the operation parameter can include, but is not limited to, at least one of the number of operations, the frequency of operations, and the time of operations.
[0038] It can be understood that when the first delivery address is modified multiple times within a preset time period (such as within 1 hour), there is a possibility that the user tries to avoid the restriction rules of the platform and realizes high-frequency ordering by transforming the address.
[0039] Step S102, determining the target credibility of the first operation based on the target parameter; the target credibility is used to judge whether the first operation is abnormal.
[0040] Here, the target credibility is a quantitative indicator that indicates whether the first operation is credible. The higher the target credibility, the less likely the first operation is abnormal. Conversely, the lower the target credibility, the more likely the first operation is abnormal.
[0041] In some embodiments, the first operation may be determined to be free of abnormality if the target credibility is not less than a credibility threshold, wherein the credibility threshold is a minimum credibility indicating that the first operation is free of abnormality, and the credibility threshold may be determined based on an empirical value or a calibration value from a pre-experiment.
[0042] During implementation, different target parameters may correspond to different credibility judgment conditions, and different credibility judgment conditions may correspond to different credibility determination methods.
[0043] For example, when the target parameters are the operation time and number of operations of the first operation, the judgment condition corresponding to the credibility can be characterized as that the user cannot modify the delivery address more than 3 times in one day. When it does not exceed 3 times, the corresponding credibility is 100%. The credibility decreases by 10% each time it exceeds the limit until it reaches 0, and the credibility threshold is 80%; based on the operation time and number of operations when the user performs the first operation, when it is judged that the user modifies the delivery address 10 times in one day, it can be determined that the target credibility corresponding to the first operation is 30%, which is less than the credibility threshold of 80%, and therefore it is judged that there is an abnormality in the first operation.
[0044] For another example, there can be multiple criteria for the credibility, and for each criterion, a target sub-credibility can be determined. Furthermore, the final target credibility of the first operation can be determined based on these multiple sub-credibilities. This allows for combining multiple criteria to improve the accuracy of the target credibility, thereby reducing the possibility of misjudging the first operation as abnormal or non-abnormal due to inaccurate target credibility determined based on a single criterion.
[0045] In some implementations, when an address change is detected, whether the address change operation is credible can be determined; or after the address change, when it is detected that the user places an order with the new address, whether the address change operation is credible can be determined.
[0046] For example, upon detecting a first operation, an operational risk assessment mechanism can be triggered to determine the target credibility of the first operation based on the target parameters corresponding to the first operation. This allows for real-time determination of whether the user's address change operation is abnormal after responding to a user's address change, and for prompts or warnings to be issued, reducing the amount of data processing required when placing an order and improving processing efficiency.
[0047] For example, when the first operation is detected and the user places an order based on the second delivery address obtained after the first operation, the operation risk assessment mechanism can be triggered. Based on the target parameter corresponding to the first operation, the target credibility of the first operation is determined. In this way, in the case where the user does not place an order based on the second delivery address, it can be considered that there is no risk for the time being, and therefore there is no need to perform operation risk assessment after adjusting the delivery address each time, thereby reducing the overall data processing amount and reducing the operating cost of the platform.
[0048] In the embodiments of the present application, in response to the editing operation of the first delivery address, i.e., the first operation, the target parameter corresponding to the first operation is obtained; based on the target parameter, the target credibility of the first operation is determined to judge whether the first operation is abnormal. In this way, in response to the editing operation of the user's delivery address, the credibility of the address editing operation can be judged in advance according to the editing information representing the change of the original delivery address and / or the operation parameter representing the editing operation behavior information before delivery to the new delivery address edited, and it is determined whether the editing operation is abnormal, thereby reducing the possibility of delivery to the abnormal address or risk address corresponding to the first operation, and further improving the safety and stability of the e-commerce platform.
[0049] In some embodiments, the above step S102 can include the following steps S111 to S112:
[0050] Step S111, based on the target parameter, determining a target feature between the first delivery address and the second delivery address; the second delivery address is obtained after performing the editing operation on the first delivery address.
[0051] Here, the first delivery address and the second delivery address are two delivery address text information submitted by the user at different time periods, and the submission time of the second delivery address is later than that of the first delivery address.
[0052] The target feature is used to represent the difference between the first delivery address and the second delivery address. Based on the target feature, the credibility of the first operation is determined, i.e., according to the difference between the first delivery address and the second delivery address, it is judged whether the second delivery address obtained after editing the first delivery address is abnormal.
[0053] Step S112, based on the target feature, determining the target credibility of the first operation.
[0054] Among them, the target feature includes a first feature, a second feature, and / or a third feature, the first feature is used to represent the matching degree of the first delivery address and the second delivery address in the semantic dimension and the spatial dimension, the second feature is used to describe the text change process between the first delivery address and the second delivery address, and the third feature is used to represent the operation frequency of the editing operation.
[0055] Here, the semantic dimension can be used to determine the similarity between the texts of the consignee addresses, and the spatial dimension can be used to determine the similarity of the consignee addresses in geographical space. According to the first feature, determining the matching degree of the first consignee address and the second consignee address in the semantic dimension and the spatial dimension can reduce the possibility of misjudging that the first operation is abnormal or non-abnormal due to inaccurate determination of the target credibility according to a single dimension similarity.
[0056] According to the second feature, determining the text change process between the first consignee address and the second consignee address can identify whether there is a regular mechanical tampering behavior or an abnormal text change in the first operation, so as to determine the credibility of the first operation and the effectiveness of the second consignee address from the dimension of text change.
[0057] The third feature of representing the operation frequency of the editing operation can reflect the change activity degree of the first consignee address. According to the third feature, potential abnormal operation can be identified in the case that the user edits the first consignee address too actively in a short time.
[0058] It can be understood that according to the above first feature, the second feature, and the third feature, the address change mode between the first consignee address and the second consignee address can be more comprehensively described, so that the target credibility of the first operation determined can more accurately evaluate the credibility of the first operation, and further improve the accuracy of abnormal address change identification.
[0059] In some embodiments, the first feature can be determined based on the text similarity and the spatial similarity between the first consignee address and the second consignee address.
[0060] In some embodiments, the second feature can be determined based on the text change content between the first consignee address and the second consignee address.
[0061] In some embodiments, the third feature can be determined based on the operation time and the operation frequency of the first operation. The third feature can be determined by detecting the operation time of the first operation in real time.
[0062] It can be understood that under normal circumstances, the consignee address of the user is relatively fixed. Considering that the user will actively modify the consignee address in the case of order failure or possible order failure, the possibility of the user having a strong motivation to bypass the risk control limit is higher in the case of high-frequency active modification.
[0063] In some embodiments, the operation parameter of the first operation within a preset time length can be counted to determine the frequency detection score of the first operation; in the case that the number of times of the first operation within the preset time length is multiple, each editing operation can correspond to a sub-frequency detection score, and the frequency detection score of the first operation within the preset time length can be determined according to the sub-frequency detection scores corresponding to all editing operations within the preset time length.
[0064] Exemplarily, the time sequence of the first operation corresponding to the first delivery address change event can be acquired. Wherein, the data corresponding to each first delivery address change event in the time sequence includes the change time of the first delivery address change event, and the time difference Δt between the last change; the time difference Δt is used to calculate the frequency detection score;
[0065] Further, a time window attenuation factor τ can be added as a time attenuation coefficient to improve the influence of the address change operation within the nearest target time length (such as 1 hour, half an hour, etc.), so as to improve the weight of the address change event within the nearest time. For example, τ can be set to 3600 to improve the weight of the address change event within the nearest hour.
[0066] Exemplarily, the process of calculating the frequency detection score score of the first operation corresponding to the first operation can refer to formula (1): freq
[0067]
[0068] Wherein, Δt i is the time difference between the i-th editing operation and the i-1-th editing operation, N is the total number of editing operations corresponding to the first operation, and e is the natural logarithm.
[0069] In the embodiments of the present application, the target feature between the first delivery address and the second delivery address obtained after the editing operation on the first delivery address is determined based on the target parameter, and the target credibility of the first operation is further determined based on the target feature; wherein, the target feature includes the matching feature of the semantics and space between the first delivery address and the second delivery address, the text change process feature, and / or the operation frequency feature of the editing operation. In this way, the target credibility of the first operation can be judged based on one or more of the multiple dimension features, the range of identifying the address abnormal change event is improved, the possibility of delivering the abnormal address or the risk address corresponding to the first operation is reduced, and the safety and stability of the e-commerce platform are further improved.
[0070] In some embodiments, the above step S111 can include the following steps S121 to S123:
[0071] In step S121, a target semantic difference degree between the first delivery address and the second delivery address is determined based on the target parameter.
[0072] Here, the target semantic difference degree can represent a substantial difference degree of the first delivery address and the second delivery address in text expression.
[0073] For example, in the case that the first delivery address and the second delivery address only differ in the house number, the target semantic difference degree is low, and it can be considered that the substantial difference degree of the first delivery address and the second delivery address in text expression is small, and the first delivery address and the second delivery address are similar; in the case that the province, city and district information of the first delivery address and the second delivery address are greatly changed, the target semantic difference degree is high, and it can be considered that the substantial difference degree of the first delivery address and the second delivery address in text expression is large, and the first delivery address and the second delivery address are not similar.
[0074] By introducing the target semantic difference degree, the change mode of the first delivery address and the second delivery address in the text level can be effectively identified, so as to find the potential abnormal change behavior, for example, the behavior of frequently modifying the house number to evade the risk control rules.
[0075] In some embodiments, the target parameter corresponding to the target semantic difference degree between the first delivery address and the second delivery address can be information related to the text corresponding to the first delivery address and the second delivery address. For example, it can include but is not limited to at least one of the first text corresponding to the first delivery address, the second text corresponding to the second delivery address, the first text vector corresponding to the first delivery address, the second text vector corresponding to the second delivery address, the text edit distance between the first delivery address and the second delivery address, etc.
[0076] In step S122, a target spatial similarity between the first delivery address and the second delivery address is determined based on the target parameter.
[0077] The target spatial similarity can represent the similarity degree of the first delivery address and the second delivery address in geographical space. The higher the target spatial similarity is, the closer the actual geographical positions of the first delivery address and the second delivery address are; on the contrary, the lower the target spatial similarity is, the farther the actual geographical positions of the first delivery address and the second delivery address are.
[0078] In some embodiments, the geographical coordinates of the address text can be obtained by parsing the original address text input by the user, by calling the application interface of the map application, POI matching, etc., to calculate the target spatial similarity.
[0079] In implementation, after obtaining the geographic coordinates of the address text, the address text can be converted into standard geographic coordinates, and the target spatial similarity is further calculated. For example, after obtaining the plane coordinates corresponding to the address text, the plane coordinates are converted into latitude and longitude coordinates.
[0080] By introducing the target spatial similarity, the rationality of the address change operation can be verified at the physical space level, the risk caused by abnormal orders due to sudden changes in geographic location can be reduced, and the accuracy of address change detection can be further improved.
[0081] In step S123, the first feature is determined based on the target semantic difference and the target spatial similarity.
[0082] Here, the first feature determined based on the target semantic difference and the target spatial similarity can be used to determine whether the semantic dimension feature and the spatial dimension feature between the first delivery address and the second delivery address conflict, that is, whether the difference in the semantic dimension and the similarity in the spatial dimension match.
[0083] For example, in the case that the target semantic difference between the first delivery address and the second delivery address is large, it indicates that the actual location difference between the first delivery address and the second delivery address is large, and the target spatial similarity should be low. If the target spatial similarity is high, the rationality of the address change is low, and the first operation may be abnormal, and the target credibility is low. If the target spatial similarity is low, the rationality of the address change is high, and the possibility of the first operation being abnormal is low, and the target credibility is high.
[0084] By combining the target semantic difference and the target spatial similarity to determine the first feature, multi-dimensional cross verification of the first operation can be realized, so that the rationality of the address change can be more comprehensively judged, and the precision and robustness of the editing operation abnormality detection can be improved.
[0085] In the embodiments of the present application, after determining the target semantic difference and the target spatial similarity between the first delivery address and the second delivery address based on the target parameters, the first feature is determined by combining the target semantic difference and the target spatial similarity. In this way, scenarios in which the text difference between the two delivery addresses before and after the editing operation does not match the actual address location change can be identified, thereby improving the accuracy of address abnormal change detection and further improving the security and stability of the e-commerce platform.
[0086] In some embodiments, the editing information of the first delivery address includes editing process information, and the editing process information includes an editing distance between the first delivery address and the second delivery address. The above step S121 can include the following steps S131 to S133:
[0087] In step S131, the editing distance between the first delivery address and the second delivery address is determined.
[0088] Here, the edit distance is used to calculate the similarity between texts, representing the minimum number of character editing operations required to convert one string to another between two strings. Among them, character editing operations can include replacing a character, inserting a character, or deleting a character. The smaller the edit distance, the more similar the two strings.
[0089] For example, in the case where there is only one character replacement operation between the string corresponding to the first delivery address and the string corresponding to the second delivery address, the edit distance is 1.
[0090] It can be understood that the edit distance between the first delivery address and the second delivery address as the edit process information can intuitively reflect the difference in address text structure in the address editing process corresponding to the first operation, and is particularly suitable for detecting subtle modifications of fields such as house numbers and street names that are relatively fixed in position.
[0091] For example, the edit distance between the first delivery address addr1 and the second delivery address addr2 can be represented as edit distance .
[0092] Step S132, determining the vector similarity between the first text vector and the second text vector; the first text vector is a text vector corresponding to the first delivery address, and the second text vector is a text vector corresponding to the second delivery address.
[0093] Here, the text vector is a way of converting address text into a numerical structure. The vector similarity between the first text vector and the second text vector can be used to measure the closeness of the two address texts in the semantic space. The higher the vector similarity between the first text vector and the second text vector, the more similar the address texts of the first delivery address and the second delivery address in semantics.
[0094] By calculating the vector similarity between the text vectors, the similarity of the address texts can be evaluated from the semantic level, thereby making up for the insensitivity of the edit distance to syntax, expression methods, etc. Combined with the edit distance and the vector similarity, the influence of spelling errors or inconsistent abbreviations can be reduced, thereby more comprehensively and accurately determining whether the first delivery address and the second delivery address are similar or have substantial differences.
[0095] For example, the process of determining the vector similarity vec addr1 between the first text vector V addr2 corresponding to the first delivery address and the second text vector V sim corresponding to the second delivery address can refer to formula (2):
[0096] vecsim = cosine similarity (V addr1 , V addr2 ) (2);
[0097] wherein cosine similarity represents calculating the cosine similarity between two vectors.
[0098] Step S133, determining the target semantic difference degree based on the edit distance and the vector similarity.
[0099] Here, the target semantic difference degree is the final evaluation result obtained by combining the edit distance and the vector similarity, which is used to quantify the overall difference between the first consignee address and the second consignee address in the text content and semantic level.
[0100] By combining the edit distance and the vector similarity, the substantive impact of address change can be more accurately evaluated. For example, even if the first consignee address and the second consignee address are slightly different in literal, but if the semantic similarity is high, it can still be considered that the first consignee address and the second consignee address point to the same geographical location; on the contrary, if the semantic difference is large, it means that the geographical locations pointed by the first consignee address and the second consignee address respectively are quite different. Combining the edit distance and the vector similarity helps to identify the behavior of circumventing the rules through minor address text modification.
[0101] In some embodiments, the edit distance can be normalized to obtain the text edit similarity between the first consignee address and the second consignee address, and then the text edit similarity and the vector similarity are weighted and summed to obtain the target semantic similarity; further, the target semantic difference degree is determined according to the target semantic similarity.
[0102] Exemplarily, the process of determining the text edit similarity edit sim based on the edit distance between the first consignee address and the second consignee address can refer to formula (3):
[0103]
[0104] wherein len(addr1) and len(addr2) are the string lengths corresponding to the texts of the first consignee address and the second consignee address respectively, max(len(addr1), len(addr2)) represents the maximum value of the string lengths corresponding to the first consignee address and the second consignee address respectively, and by taking the maximum length of the two addresses as the denominator, the edit distance can be normalized, so that the value range of the text edit similarity edit sim is between 0 and 1.
[0105] Exemplarily, the text editing similarity edit between the first consignee address and the second consignee address is calculated according to the following formula (2): sim And the vector similarity vec between the first consignee address and the second consignee address is calculated according to the following formula (3): sim After weighting, the target semantic similarity text is obtained by summing up the text editing similarity edit and the vector similarity vec. sim The process can be seen from formula (4):
[0106] text = a * vec + b * edit (4); sim sim sim
[0107] Wherein, a and b are weights of the vector similarity vec and the text editing similarity edit between the first consignee address and the second consignee address, respectively. The greater the target semantic similarity text is, the closer the contents of the first consignee address and the second consignee address are. sim sim sim
[0108] Exemplarily, the values of a and b can be 0.6 and 0.4, respectively.
[0109] In the case of the target semantic similarity text, the target semantic difference degree can be represented as 1-text. sim sim
[0110] In the embodiment of the application, by respectively determining the double text similarity evaluation features (editing distance and vector similarity) between the first consignee address and the second consignee address, the target semantic difference degree between the first consignee address and the second consignee address is determined. In this way, the rationality of address change can be evaluated from both character level and semantic level, the error of the target text difference degree is reduced, thereby improving the accuracy of abnormal address change detection, and further improving the accuracy of identifying abnormal address change behavior.
[0111] In some embodiments, the editing information of the first consignee address includes editing result information, and the editing result information includes the second consignee address obtained by editing the first consignee address. The above step S122 can include the following steps S141 to S142:
[0112] Step S141, based on the geographic location information corresponding to the first consignee address and the second consignee address respectively, the spherical distance between the first consignee address and the second consignee address is determined.
[0113] Here, the second consignee address is the editing result information, that is, the editing result corresponding to the first operation, which is used for similarity comparison with the first consignee address, so as to determine the address information change difference before and after the first operation is performed.
[0114] The geographic position information is information representing the specific position of an address in a geographic space, and may include, for example, longitude and latitude coordinates, geodetic coordinates, etc.
[0115] In some embodiments, after obtaining the longitude and latitude information of the first and second delivery addresses, the spherical distance can be calculated according to the Haversine formula.
[0116] For example, the process of determining the spherical distance Haversine distance between the first and second delivery addresses can refer to formula (5):
[0117] distance = Haversine distance = (Geo1(lon1, lat1), Geo2(lon2, lat2)) (5);
[0118] wherein Geo1(lon1, lat1) is the longitude and latitude coordinates of the POI point Geo1 corresponding to the first delivery address, Geo2(lon2, lat2) is the longitude and latitude coordinates of the POI point Geo2 corresponding to the second delivery address, and distance is the spherical distance Haversine distance .
[0119] Step S142, determining the target spatial similarity based on the spherical distance and a preset distance baseline value; the distance baseline value is used to determine the proximity of the first and second delivery addresses in the geographic space.
[0120] Here, the distance baseline value can be an empirical value or a self-defined calibration value, which is used to measure whether the first and second delivery addresses are in the same range in the geographic space. For example, in the case that the spherical distance between the first and second delivery addresses is not greater than the distance baseline value, it can be considered that the two addresses are in the same adjacent area; otherwise, it can be considered that the first and second delivery addresses have significant differences in the geographic space.
[0121] In some embodiments, the distance baseline value can be dynamically adjusted according to the business scenario to adapt to the needs of different regions and / or different risk levels.
[0122] For example, the distance baseline value can be set to 250, indicating that in the case that the spherical distance between the first and second delivery addresses is not greater than 250 meters, it is considered that the two addresses are in the same adjacent area and are relatively close; in the area with large land area and small population density, considering the problem of large land area and far distance between buildings, the distance baseline value can be increased to 300.
[0123] Exemplarily, the target spatial similarity geo between the first consignee address and the second consignee address is determined. sim The process can be seen from formula (6):
[0124]
[0125] Wherein, μ is the distance baseline value; according to formula (6), the greater the spherical distance distance between the first consignee address and the second consignee address, the smaller the target spatial similarity.
[0126] Exemplarily, referring to formula (4) to (6), based on the target semantic difference 1-text sim and the target spatial similarity geo sim between the first consignee address and the second consignee address, the process of determining the first feature value score conflict corresponding to the first feature can be seen from formula (7):
[0127]
[0128] Wherein, the first feature value score conflict is negatively related to the target semantic similarity text sim , and is positively related to the target spatial similarity geo sim , the greater the first feature value score conflict , the greater the conflict between the text difference and the geographic coordinates between the first consignee address and the second consignee address; therefore, the first feature value score conflict can be used to identify the scene that the text difference is large but the actual geographic position is close, so as to reduce the possibility of avoiding risk control rules only through text difference, and further improve the accuracy of address abnormal change detection.
[0129] In the embodiment of the application, the spherical distance between the first consignee address and the second consignee address is calculated based on the geographic position information, and the target spatial similarity between the first consignee address and the second consignee address is determined in combination with the preset distance baseline value. In this way, the spherical distance between the first consignee address and the second consignee address and the preset baseline value can be used to improve the accuracy of evaluating the target spatial similarity between the first consignee address and the second consignee address.
[0130] In some embodiments, the above step S111 can include the following steps S151 to S153:
[0131] Step S151, based on the target parameter, determining the character change feature between the first text corresponding to the first consignee address and the second text corresponding to the second consignee address.
[0132] Here, the character change feature can represent the character-level difference between the first text and the second text, and can be used to describe the character editing actions and their distribution in the address change process corresponding to the first operation.
[0133] The character change feature can be used to identify abnormal editing patterns in the address change process, such as frequent modifications to end information such as house numbers. This helps to detect mechanical dynamic adjustments to the first delivery address, further improving the accuracy of address change anomaly identification.
[0134] For example, the character change feature can be used to represent the number of different character editing actions, whether the character change position is concentrated, etc.
[0135] In some embodiments, the target parameters for determining the character change feature between the first text corresponding to the first delivery address and the second text corresponding to the second delivery address can include, but are not limited to, at least one of the character type, the number of characters, the character position, etc. that are changed in the second text relative to the first text.
[0136] Step S152, based on the target parameters, determines the administrative region change feature between the first text corresponding to the first delivery address and the second text corresponding to the second delivery address.
[0137] Here, the administrative region change feature can be used to represent whether the path change of the administrative region is reasonable when changing from the first delivery address to the second delivery address, i.e., whether the change path of the first delivery address and the second delivery address on the administrative level conforms to the conventional logic.
[0138] For example, changing the delivery address from "Beijing Haidian District Zhongguancun Street" to "Beijing Chaoyang District Wangjing Street" can be considered as a normal administrative region jump; but if the delivery address is changed from "Sichuan Province Chengdu Wuhou District" to "Hubei Province Wuhan Jianghan District", it can be considered that the administrative region jump path is irregular, and the new delivery address has a high risk.
[0139] It can be understood that based on the administrative region change feature, the unnatural jump path in the address change process can be identified, thereby capturing the behavior of bypassing the risk control limit by changing the address across regions, further improving the identification ability of abnormal addresses.
[0140] In some embodiments, the target parameters for determining the administrative region change feature between the first text corresponding to the first delivery address and the second text corresponding to the second delivery address can include, but are not limited to, the second text corresponding to the second delivery address, the administrative region information in the second text, the administrative region information in the second text that is changed relative to the first text, etc.
[0141] Step S153, determining the second feature based on the character change feature and the administrative region change feature.
[0142] Here, the second feature is a comprehensive feature after fusing the character change feature and the administrative region change feature, so that the risk level of the address change can be evaluated from multiple dimensions at the character level and the administrative region level, and the accuracy is higher.
[0143] In the embodiments of the present application, the character change feature and the administrative region change feature are introduced and fused as the second feature for determining the credibility of the first operation target. In this way, the accuracy of determining whether the first operation is abnormal using the second feature can be improved according to the character change in the first text and the second text and the change of the corresponding administrative region.
[0144] In some embodiments, the editing operation includes an editing action sequence, the editing action includes an adding action, a deleting action, or a modifying action, the operation parameter of the first operation includes a target proportion of the modifying action in the editing action sequence, and the editing information of the first consignee address includes editing process information, the editing process information includes a number of target characters modified in the first text and positions of the target characters in the first text. The above step S151 can include the following steps S161 to S162:
[0145] Step S161, determining the modifying concentration of the first text based on the position identification value of each target character in the first text, the mean value of the position identification values of all target characters, and the number of target characters.
[0146] Here, by analyzing the editing action sequence corresponding to the editing operation, it can be detected whether the address text contains mechanical tampering features in the continuous editing actions. In the case of containing mechanical tampering features, the first operation corresponding to the editing action sequence is likely to be a corresponding high-risk abnormal operation.
[0147] In some embodiments, the editing action sequence in which the complete address changes twice in succession can be extracted as the editing action sequence corresponding to an editing operation.
[0148] The position identification value is used to identify the position of the target character modified in the first text in the original character string.
[0149] For example, the position identification value of each target character in the first text can be the number corresponding to the target character in sequence after numbering the characters in the first text from 0.
[0150] For example, for the first text with the text content of "Chaoyang District Guangshunbei Street No. 16" and the second text with the text content of "Chaoyang District Guangshunbei Street No. 17", the target character "6" with the position identification value of 9 is modified.
[0151] In some embodiments, in the case that the first operation corresponds to an editing action sequence including editing actions of different types, the editing action sequence can be divided into a plurality of sub-editing action sequences corresponding to the editing action types according to the editing action types, the starting positions and the ending positions of the target characters corresponding to each editing action type in the first text, and the starting positions and the ending positions of the target characters in the second text.
[0152] For example, an adding action in the editing action can be represented as insert, a deleting action can be represented as delete, and a modifying action can be represented as replace. In addition, the same characters at the same positions in the first text and the second text can be represented as equal.
[0153] That is, for the example above, the first text with the text content of "Chaoyang District Guangshun North Street No. 16" and the second text with the text content of "Chaoyang District Guangshun North Street No. 17", the editing action sequence corresponding to the first operation can be represented as [(“equal”, 0, 9, 0, 9), (“replace”, 9, 10, 9, 10)].
[0154] The first sub-editing action sequence (“equal”, 0, 9, 0, 9) indicates that the characters “Chaoyang District Guangshun North Street 1” corresponding to the numbers 0 to 8 in the first text are completely identical to the characters “Chaoyang District Guangshun North Street 1” corresponding to the numbers 0 to 8 in the second text, and the change starts from the character corresponding to the number 9 in the first text. The second sub-editing action sequence (“replace”, 9, 10, 9, 10) indicates that the character “6” corresponding to the number 9 in the first text is replaced by the character “7” corresponding to the number 9 in the second text, and the replacing action ends before the character “No.” corresponding to the number 10. The number of target characters is 1, and the position identifier value corresponding to the target character is 9.
[0155] The modification concentration of the first text is used to indicate whether the distribution of the target characters is concentrated.
[0156] In some embodiments, the modification concentration of the first text can be determined based on the mean value of the position identifier values of all target characters.
[0157] For example, after obtaining the position identifier value of each target character and the number of all target characters, the mean value of the position identifier values of all target characters can be calculated, and the absolute value of the difference between the mean value and the median of all position identifier values can be determined. The reciprocal of the absolute value is taken as the modification concentration of the first text. The greater the modification concentration, the more concentrated the distribution of the target characters. The smaller the modification concentration, the more dispersed the distribution of the target characters.
[0158] In some embodiments, the modification concentration of the first text can also be determined based on a variance of the number of the target character in the first text.
[0159] For example, the process of determining the variance of the number of the target character in the first text can refer to formula (8):
[0160]
[0161] wherein the variance σ of the number of the target character in the first text 2 corresponding to the modification concentration of the first text, M is the total number of the target character in the editing action sequence in which the adjacent two complete addresses are changed, x j is the position identification value of the target character of the jth replacement, is the average value of the position identification values of all target characters.
[0162] For example, in the case that the modification action is performed 4 times in the editing action sequence corresponding to the first operation, the position identification values of the target character in the first text are 9, 10, 11 and 9 respectively, and the average value of the position identification values is 9.75. 2 The total variance σ 2 is about 0.6875, which is obtained by [(-0.75) 2 , 0.25 2 , 1.25 2 , (-0.75) 2 ] / 4.
[0163] For example, 1 / (1+σ 2 ) can be used as the modification concentration of the first text. The larger the variance is, the smaller the modification concentration is, which indicates that the distribution of the characters corresponding to the modification action is more dispersed. The smaller the variance is, the larger the modification concentration is, which indicates that the distribution of the characters corresponding to the modification action is more concentrated.
[0164] By determining the modification concentration of the first text, the intensive replacement operation of the fixed character can be identified, and the accuracy of the abnormal address identification is further improved.
[0165] In step S162, the character change feature is determined based on the target proportion and the modification concentration.
[0166] Here, the target proportion is the ratio of the number of the modification action to the total number of the editing action in the editing action sequence corresponding to the first operation.
[0167] In some embodiments, the target proportion can be determined after the number of each editing action in the editing action sequence corresponding to the first operation is counted.
[0168] For example, the target proportion Ratio replaceThe process of determining the character change feature can refer to formula (9):
[0169]
[0170] wherein Count replace is the number of modified actions, Count insert is the number of added actions, and Count delete is the number of deleted actions.
[0171] It can be understood that the higher the target proportion, the more the user tends to modify the address by replacing characters rather than by adding or deleting characters, and in this case, the user is more likely to evade the risk control rules.
[0172] The character change feature determined in combination with the target proportion and the modification concentration can describe the overall trend of the editing action corresponding to the first operation, and can be used to identify character changes that occur only at the end or in a concentrated position in a continuously modified delivery address.
[0173] In some embodiments, the ratio of the target proportion to the modification concentration can be determined as the character change feature score seq .
[0174] For example, in the case where the modification concentration of the first text is 1 / (1+σ 2 ), the process of determining the character change feature can refer to formula (10):
[0175]
[0176] It can be understood that the higher the target proportion Ratio replace , or the larger the variance σ 2 , that is, the smaller the modification concentration 1+σ 2 , the more concentrated the change position, the more likely the first operation is an abnormal operation, and the greater the value of the corresponding character change feature score seq .
[0177] For example, in the case where the target proportion of the modification operation is low (for example, Ratio replace <20%), or the position of the target character is relatively dispersed and the position variance is large, the value of the character change feature score seq is small, and the first operation can be considered a normal operation.
[0178] For another example, in the case where the target proportion of the modification operation is high (for example, Ratio replace >60%), or the position of the target character is relatively concentrated (such as frequently modifying the house number), the position variance is small, and can even be 0, the value of the character change feature scoreseq The value of is large, and the first operation can be regarded as an abnormal operation.
[0179] In this embodiment, the character change characteristics are further determined by introducing position identification values, modification concentration, and target proportion. This allows for a more comprehensive assessment of the editing operations during the address change process based on the type of editing action corresponding to the first operation, thereby improving the accuracy of identifying mechanical character modifications and further improving the accuracy of detecting abnormal address changes.
[0180] In some embodiments, the edit information of the first delivery address includes edit result information, and the edit result information includes administrative area information of the second delivery address obtained after editing the first delivery address. The above step S152 includes the following steps S171 to S172:
[0181] Step S171: Determine, from the second text, second administrative region information that is different from first administrative region information of the same administrative region level in the first text.
[0182] Here, administrative regions are defined as multiple hierarchical units within an address structure, divided by administrative management. Each administrative region has a subordinate relationship with another, forming a top-down administrative system. For example, administrative regions may include, but are not limited to, provinces, cities, districts, counties, and sub-districts, from top to bottom.
[0183] Administrative region information is specific information corresponding to the administrative region level, for example, it may include but is not limited to province information, city information, district and county information, and street information.
[0184] For example, for the address "Guixi Street, Wuhou District, Chengdu City, Sichuan Province", the administrative region information corresponding to the provincial administrative region level is Sichuan Province, the administrative region information corresponding to the municipal administrative region level is Chengdu City, the administrative region information corresponding to the district and county administrative region level is Wuhou District, and the administrative region information corresponding to the street-level administrative region level is Guixi Street.
[0185] The first and second administrative region information correspond to the same administrative region level, but the specific information is different. When the first and second texts contain multiple administrative region levels with different administrative region information, the first and second administrative region information can be all different administrative region information at the same administrative region level.
[0186] For example, when the first text corresponding to the first delivery address is "Guixi Street, Wuhou District, Chengdu City, Sichuan Province", and the second text corresponding to the second delivery address is "Xihua Street, Fucheng District, Mianyang City, Sichuan Province", the first administrative area information is "Guixi Street, Wuhou District, Chengdu City", and the second administrative area information is "Xihua Street, Fucheng District, Mianyang City".
[0187] In some embodiments, after obtaining the delivery address, the address can be standardized at the administrative region level.
[0188] In some embodiments, the address can be divided based on a plurality of preset administrative region levels.
[0189] For example, the administrative region levels can be set from high to low as four levels: S1 corresponds to the province level, S2 corresponds to the city level, S3 corresponds to the county level, and S4 corresponds to the street level.
[0190] In implementation, the administrative region levels corresponding to the address containing a municipality such as Beijing can be specially processed, for example, the values of the region levels S1 and S2 can be determined as the values corresponding to the municipality; for example, the address of “Beijing Haidian District Shangdi Street” can be divided according to the administrative region levels as: S1 (Beijing), S2 (Beijing), S3 (Haidian District), and S4 (Shangdi Street).
[0191] Step S172, based on the administrative region transition probability corresponding to the administrative region information of each administrative region level in the second administrative region information, determine the administrative region change feature; for the administrative region information of each administrative region level, the corresponding administrative region transition probability is the probability of transition from the administrative region information of the previous level of the administrative region level to the administrative region information of the administrative region level.
[0192] Here, the administrative region transition probability is used to represent the possibility of transition from the administrative region information corresponding to one administrative region level to the administrative region information corresponding to another administrative region level.
[0193] For example, for the municipal administrative region information “Mianyang City”, the corresponding administrative region transition probability is the transition probability from the provincial administrative region information “Sichuan Province” to the municipal administrative region information “Mianyang City”.
[0194] For example, for the county-level administrative region information “Fucheng District”, the corresponding administrative region transition probability is the transition probability from the municipal administrative region information “Mianyang City” to the county-level administrative region information “Fucheng District”.
[0195] In some embodiments, by constructing a multi-order Markov model, the state transition probability of the corresponding administrative region information between different administrative region levels can be calculated, so as to determine the administrative region change feature between the first delivery address and the second delivery address, and further determine whether the administrative region path changed in the second delivery address relative to the first delivery address is abnormal (i.e. in the case of too low state transition probability, the path transition from the first delivery address to the second delivery address is a low probability event).
[0196] Exemplarily, in a case where the first text corresponding to the first consignee address is "Guixi Street, Wuhou District, Chengdu City, Sichuan Province", and the second text corresponding to the second consignee address is "Xihua Street, Fucheng District, Mianyang City, Sichuan Province", the administrative region information corresponding to the S2 level in the first text and the second text is "Sichuan Province"; the administrative region information corresponding to the S3 level in the first text is "Wuhou District", which is different from "Fucheng District" in the second text; the administrative region information corresponding to the S4 level in the first text is "Guixi Street", which is different from "Xihua Street" in the second text; therefore, the first administrative region information is "Guixi Street, Wuhou District, Chengdu City, Sichuan Province", and the second administrative region information is "Xihua Street, Fucheng District, Mianyang City, Sichuan Province";
[0197] In this case, the administrative region transition probability P path corresponding to the second administrative region information "Xihua Street, Fucheng District, Mianyang City, Sichuan Province" is determined to be 0.5.
[0198] P path = P (绵阳市|四川省) * P (涪城区|绵阳市) * P (西华街道|涪城区) (11).
[0199] Wherein, P (绵阳市∣四川省) represents the state transition probability corresponding to the municipal administrative region level is the probability of transition from Sichuan Province to Mianyang City, P (涪城区∣绵阳市) represents the state transition probability corresponding to the county-level administrative region level is the probability of transition from Mianyang City to Fucheng District, and P (西华街道∣涪城区) represents the state transition probability corresponding to the street-level administrative region level is the probability of transition from Fucheng District to Xihua Street.
[0200] It can be understood that for the path corresponding to the administrative region level, the greater the change of the path in the administrative region dimension, the greater the possibility of abnormal address change, and the over transition of the corresponding user space state, at this time, the rationality of the path transition needs to be further determined according to the administrative region change characteristics.
[0201] Exemplarily, based on the administrative region transition probability P path corresponding to the second administrative region information "Xihua Street, Fucheng District, Mianyang City, Sichuan Province", the process of determining the administrative region change characteristics can be referred to formula (12):
[0202]
[0203] Wherein, 10 -∈ is a small constant, which avoids the abnormal calculation result in the case that P path is 0, and the smaller the P path , the higher the score markovThe higher, the higher the possibility that the second text corresponds to an abnormal second delivery address.
[0204] In the embodiments of the present application, the administrative region change feature is determined based on the administrative region transition probability corresponding to the administrative region information of each administrative region level in the second administrative region information. In this way, the abnormal jump of the change path can be identified through the state transition probability of the administrative region, thereby further improving the accuracy of the abnormal address change detection.
[0205] In some embodiments, the above processing method can further include the following steps S181 to S182:
[0206] Step S181, obtaining a target address data set; the target address data set includes a target number of target delivery addresses.
[0207] Here, the target address data set is a collectable address data set associated with the application scenario of the current delivery address, which can include but is not limited to the standardized address records extracted from the platform's entire historical orders, user registration information and / or logistics system.
[0208] The target number is at least 1.
[0209] The target delivery address is not limited to whether there is a corresponding order record, and is not limited to all delivery addresses of the user.
[0210] Step S182, for the administrative region information of each first administrative region level, determining the ratio of the first number of the first level address to the second number of the target upper level address as the administrative region transition probability of the administrative region information of the first administrative region level; the first level address is the target delivery address in the target address data set containing the administrative region information of the first administrative region level, and the target upper level address is the target delivery address containing the administrative region information of the first administrative region level corresponding to the administrative region information of the upper administrative region level.
[0211] Here, for each first administrative region level, the corresponding first level address is a complete address containing the administrative region information corresponding to the first administrative region level. By introducing the ratio of the first number of the first level address to the second number of the target upper level address, the weight of the administrative region information corresponding to each administrative region level in the complete address system can be quantified, thereby helping to identify the delivery address deviating from the regular path transition.
[0212] The first administrative region level is at least one administrative region level in the non-highest administrative region level, and each first administrative region level corresponds to at least one superior administrative region.
[0213] In some embodiments, the second administrative region level can be one of a municipal level, a district level, and a street level; or can be any multiple of the municipal level, the district level, and the street level except the highest provincial level.
[0214] For example, the second administrative region level can be a municipal level, and the upper level thereof can be a provincial level.
[0215] For another example, the second administrative region level can be a combination level composed of a municipal level, a district level, and a street level, and the upper level thereof can be a provincial level.
[0216] Exemplarily, the administrative region information of Chengdu City as a municipal level (which can correspond to the above administrative region level S2) can be represented as S2 (成都市) Further, a first quantity count (四川省→成都市) of target delivery addresses (i.e., first level addresses) in the target address data set containing Chengdu City is determined. (四川省→成都市) Since the upper administrative region level of Chengdu City is a provincial level (which can correspond to the above administrative region level S1), the corresponding administrative region information is Sichuan Province, the ratio of the first quantity count (四川省总样本) to a second quantity count (四川省) , i.e., the proportion of users who select “Sichuan Province” also select “Chengdu City” in the target delivery address, is determined as the transition probability P (成都市) from S1 (S2|S1) to S2 .
[0217] Similarly, the administrative region information of Wuhou District as a district level (which can correspond to the above administrative region level S3) can be represented as S3 (武侯区) Further, a first quantity count (成都市→武侯区) of target delivery addresses (i.e., first level addresses) in the target address data set containing Wuhou District is determined. (成都市→武侯区) Since the upper administrative region level of Wuhou District is a municipal level (which can correspond to the above administrative region level S2), the corresponding administrative region information is Chengdu City, the ratio of the first quantity count (成都市总样本) to a second quantity count (成都市) , i.e., the proportion of users who select “Chengdu City” also select “Wuhou District” in the target delivery address, is determined as the transition probability P (武侯区) from S2 (S3|S2) to S3 .
[0218] Similarly, the administrative region information of Guixi Street as a street level (which can correspond to the above administrative region level S4) can be represented as S4 ( 桂溪街道 ) Further, a first quantity count (武侯区→桂溪街道) of target delivery addresses (i.e., first level addresses) in the target address data set containing Guixi Street is determined.; since the last administrative region level of Guixi Street is county level (which can correspond to the above-mentioned administrative region level S3), the corresponding administrative region information is Wuhou District, therefore, the ratio of the first quantity count (武侯区→桂溪街道) and the second quantity count (武侯区总样本) is determined as the transition probability P (武侯区) from S3 (桂溪街道) to S4 (S4|S3) .
[0219] In the embodiments of the present application, by obtaining the target address data set, the target quantity of target delivery addresses in the target address data set is used to determine the transition probability between the administrative region information of different first administrative region levels. In this way, the transition between the first administrative region levels in the address change process from the first text to the second text can be determined through the administrative transition probability, so as to identify abnormal address change events that do not conform to the regular path, and further improve the accuracy of abnormal address change detection.
[0220] In some embodiments, the above processing method can further include the following step S191:
[0221] Step S191, for the administrative region information of each second administrative region level, the ratio of the third quantity of the second-level address corresponding to the target quantity is determined as the administrative region transition probability of the administrative region information of the second administrative region level; the second-level address is the target delivery address in the target address data set which contains the administrative region information of the second administrative region level; the second administrative region level is the highest administrative region level, and the second administrative region level is higher than the first administrative region level.
[0222] Here, the second administrative region level is the highest level in the delivery address; for example, the second administrative region level can correspond to the above-mentioned S1 provincial level; for another example, the second administrative region level can correspond to the above-mentioned S2 municipal level (corresponding to the case that the municipality is the highest region level).
[0223] The administrative region information corresponding to the second administrative region level is the highest level administrative region information, for example, the second administrative region level can include but is not limited to Sichuan Province, Shaanxi Province, Beijing, etc.
[0224] Exemplarily, Sichuan Province as the administrative region information of the provincial level (which can correspond to the above-mentioned administrative region level S1) can be represented as S1 (四川省) , further, the third quantity count (四川省) of the target delivery address (i.e. the second-level address) containing Sichuan Province in the target address data set is determined, and further the ratio of the third quantity count (四川省) and the target quantity count (总样本数量) is determined as the transition probability from S1 (四川省)The corresponding transition probability P (S1) .
[0225] For example, referring to the above example and formula (11), the overall transition path probability P path1 from Sichuan Province to Chengdu City, and further from Wuhou District to Guixi Street is determined as follows:
[0226] P path1 = P (S1) * P (S2|S1) * P (S3|S2) * P (S4|S3) (13).
[0227] In the embodiments of the present application, the ratio of the number of second-level addresses to the target number is determined to determine the path transition probability under the second administrative region level. In this way, the probability of the jump path from the highest administrative region level can be determined, and the accuracy of the overall transition path probability is further improved, thereby improving the accuracy of determining the administrative region features.
[0228] In some embodiments, after obtaining a plurality of features, dynamic weight distribution can be performed on the plurality of features to determine the final target credibility.
[0229] For example, after obtaining the text space difference score score conflict corresponding to the first feature F1, the character change feature value score seq and the administrative region change feature value score markov corresponding to the second feature F2 and F3, and the frequency detection score score freq corresponding to the third feature F4, the importance score Importance(F k ) of each feature is obtained by using a random forest model based on the artificially labeled training data set, and the weight corresponding to each feature is calculated based on Importance(F k ). The value of Importance(F k ) can be between 0 and 1.
[0230] For example, the process of determining the weight w k corresponding to each feature based on the importance score Importance(F k ) of each feature can be seen in formula (14):
[0231]
[0232] Further, the process of determining the risk score RiskScore based on Importance(F k ) can be seen in formula (15):
[0233]
[0234] The training can be combined with different types of business data to obtain a score threshold threshold of a risk score RiskScore under the current business type. In the case where the risk score RiskScore is higher than the score threshold threshold, it can be considered that the current address change is risky, and the first operation is abnormal.
[0235] In some embodiments, the target credibility can be inversely proportional to the risk score RiskScore. The higher the risk score, the lower the target credibility of the first operation.
[0236] In some embodiments, to adapt to the drift of data distribution, a periodic update mechanism can be designed to periodically train the model in full and update the importance and weight corresponding to each feature.
[0237] In the embodiments of the present application, risk score calculation is performed based on fused feature values, which can cover scenarios corresponding to each feature. In this way, through multi-dimensional scoring, the range of detectable scenarios can be improved, and even if some features have extreme performance, the accuracy of determining whether the first operation is abnormal can be improved by combining multiple features. For example, there may be a single feature value representing that the first operation is a normal operation in this scenario, but other feature values represent that the first operation is abnormal in the corresponding scenarios, and the abnormal risk address can also be identified.
[0238] With the launch of "trade-in" subsidy activities by the state or merchants on e-commerce platforms, there are problems of illegal profiteering by ticket scalpers, gray production, etc. For example, by changing the address to bypass the address interception rules maintained by the e-commerce platform (limiting the number of orders for the same address, limiting the number of orders for similar addresses, prohibiting orders for risky addresses, etc.); After the order is rejected, the gray production will try to change the address to avoid the above rules.
[0239] On this basis, the embodiments of the present application provide a processing method, which can be executed by an electronic device, for identifying abnormal addresses and abnormal address changes in e-commerce scenarios. As shown in Figure 2 The processing method includes the following steps S201 to S205:
[0240] Step S201, address resolution and standardization.
[0241] Here, the original address input by the user, which is an unstructured field, is converted into a standard field and geographic coordinates.
[0242] Step S202, high-frequency address change detection.
[0243] Here, the formula (1) in step S112 of the above processing method can be referred to.
[0244] In implementation, the time stamp of each address change event corresponding to the first operation can be collected, and the corresponding frequency detection score score freq .
[0245] Exemplarily, the sequence of address change events can be represented as follows:
[0246] [1717027200, #2024-05-30 00:00:00 (1st change)
[0247] 1717029000, #00:30:00 (2nd change, Δt=1800 seconds)
[0248] 1717029900, #00:45:00 (3rd change, Δt=900 seconds)
[0249] 1717030200, #00:50:00 (4th change, Δt=300 seconds)
[0250] Wherein, 1717027200, 1717029000, 1717029900, 1717030200 are the time stamps corresponding to each address change event, and the total score corresponding to the first operation is the score of the cumulative number of changes, for example, the above 2nd, 3rd and 4th changes correspond to scores 0.6065, 0.7788 and 0.9200 respectively, so the total score corresponding to the first operation is 2.3053.
[0251] Step S203, two-dimensional address similarity calculation.
[0252] Here, the two-dimensional address similarity can correspond to the first feature in the above processing method.
[0253] In implementation, the standardized address can be detected for substantial differences, i.e. text similarity calculation, to obtain the target semantic difference degree in the above processing method; further, the spherical distance between the first and second delivery addresses is calculated according to the geographic coordinates, to verify the relevance of the two geographic locations in the spatial dimension, i.e. geographic similarity calculation, to obtain the target spatial similarity in the above processing method; and based on the target semantic difference degree and the target spatial similarity, the first feature is determined.
[0254] Step S204, dynamic modeling of address change process.
[0255] In implementation, the editing action sequence can be analyzed, i.e., the character-level operation of continuous address change is recorded, including the editing action type and the character change position corresponding to the modification action in the editing action, and then the operation probability and operation variance are obtained; further, a multi-order Markov modeling is used to establish a transition probability matrix of administrative region levels, which is used to identify irregular jump paths; wherein the transition probability in the transition probability matrix can correspond to the above-mentioned administrative region transition probability.
[0256] Step S205, composite risk determination.
[0257] Here, the target credibility of the first operation in the above processing method can be determined in combination with the change frequency feature (corresponding to F4 in the above processing method) in the above step S202, the semantic-space conflict analysis (corresponding to F1 in the above processing method) in the above step S203, the editing anomaly (corresponding to F2 in the above processing method) in the above step S204, and the Markov path probability (corresponding to F3 in the above processing method), and the risk determination is comprehensively determined based on the target credibility.
[0258] In the embodiments of the present application, through address structured analysis and multi-dimensional similarity analysis, cross verification of physical space and text space is realized, and the address change process is dynamically modeled, and the evolution law of the address sequence is analyzed to capture hidden abnormal changes. In this way, on the one hand, by fusing multi-dimensional features, fusing text differences, geographical offsets, change frequencies, editing modes, transition paths and other features, the detection accuracy can be improved by breaking through the limitation of a single dimension; on the other hand, through variance analysis of the editing action sequence, regular mechanical tampering of house numbers can be detected; on the other hand, by using the Markov model of administrative region level perception, the administrative levels are distinguished, the hierarchical space is established and standardized, the probability of the modification path is calculated, and the change of abnormal low-probability administrative regions can be identified.
[0259] The embodiments of the present application provide a processing device, as shown in the figure, the processing device 300 comprises: Figure 3
[0260] The acquisition module 310 is configured to acquire a target parameter corresponding to a first operation in response to the first operation; the first operation comprises an editing operation on a first delivery address, and the target parameter comprises editing information of the first delivery address and / or operation parameters of the first operation;
[0261] The determination module 320 is configured to determine a target credibility state of a second delivery address based on the target parameter; the second delivery address is obtained after the first delivery address performs the editing operation, and the target credibility state is used to determine whether the second delivery address is valid.
[0262] In some embodiments, the determining module can include a first sub-determining module configured to: determine a target feature between the first delivery address and the second delivery address based on the target parameter; and determine a target trust status of the second delivery address based on the target feature; wherein the target feature includes a first feature, a second feature, and / or a third feature, the first feature is used to represent a matching degree between the first delivery address and the second delivery address in a semantic dimension and a spatial dimension, the second feature is used to describe a text change process between the first delivery address and the second delivery address, and the third feature is used to represent an operation frequency of the editing operation.
[0263] In some embodiments, the first sub-determining module can be further configured to: determine a target semantic difference degree between the first delivery address and the second delivery address based on the target parameter; determine a target spatial similarity between the first delivery address and the second delivery address based on the target parameter; and determine the first feature based on the target semantic difference degree and the target spatial similarity.
[0264] In some embodiments, the editing information of the first delivery address includes an edit distance between the first delivery address and the second delivery address, and the first sub-determining module can be further configured to: determine the edit distance between the first delivery address and the second delivery address; determine a vector similarity between a first text vector and a second text vector, the first text vector being a text vector corresponding to the first delivery address, and the second text vector being a text vector corresponding to the second delivery address; and determine the target semantic difference degree based on the edit distance and the vector similarity.
[0265] In some embodiments, the editing information of the first delivery address includes the second delivery address obtained by editing the first delivery address, and the first sub-determining module can be further configured to: determine a spherical distance between the first delivery address and the second delivery address based on geographic position information corresponding to the first delivery address and the second delivery address, respectively; and determine the target spatial similarity based on the spherical distance and a preset distance baseline value, the distance baseline value being used to determine a closeness degree of the first delivery address and the second delivery address in a geographic space.
[0266] In some embodiments, the first sub-determining module can be further configured to: determine, based on the target parameter, a character change feature between the first text corresponding to the first consignee address and the second text corresponding to the second consignee address; determine, based on the target parameter, an administrative region change feature between the first text corresponding to the first consignee address and the second text corresponding to the second consignee address; and determine the second feature based on the character change feature and the administrative region change feature.
[0267] In some embodiments, the editing operation includes a sequence of editing actions, and each editing action includes an adding action, a deleting action, or a modifying action. The operation parameter of the first operation includes a target proportion of the modifying action in the sequence of editing actions. The editing information of the first consignee address includes a number of target characters in the first text that are modified and positions of the target characters in the first text. The first sub-determining module can be further configured to: determine a modification concentration of the first text based on a position identifier of each target character in the first text, an average of position identifiers of all target characters, and the number of target characters; and determine the character change feature based on the target proportion and the modification concentration.
[0268] In some embodiments, the editing information of the first consignee address includes administrative region information of the second consignee address obtained after the first consignee address is edited. The first sub-determining module can be further configured to: determine, from the second text, second administrative region information that is different from first administrative region information of the same administrative region level in the first text; determine the administrative region change feature based on administrative region transition probabilities of administrative region information of each administrative region level in the second administrative region information; and for administrative region information of each administrative region level, the corresponding administrative region transition probability is a probability of transition from administrative region information of a previous level of the administrative region level to the administrative region information of the administrative region level.
[0269] In some embodiments, the determining module can include a second sub-determining module configured to: obtain a target address data set; the target address data set includes a target number of target consignee addresses; for administrative region information of each first administrative region level, determine a ratio of a first number of first-level addresses to a second number of target upper-level addresses as an administrative region transition probability of the administrative region information of the first administrative region level; the first-level address is a target consignee address in the target address data set that includes the administrative region information of the first administrative region level, and the target upper-level address is a target consignee address that includes administrative region information of a previous administrative region level corresponding to the administrative region information of the first administrative region level.
[0270] In some embodiments, the determining module can include a third sub-determining module configured to: for administrative region information of each second administrative region level, determine, as an administrative region transfer probability of the administrative region information of the second administrative region level, a ratio of a third number of second-level addresses to the target number, the second-level addresses being target delivery addresses in the target address data set containing the administrative region information of the second administrative region level, the second administrative region level being the highest administrative region level, and the second administrative region level being higher than the first administrative region level.
[0271] An electronic device is provided, including a memory and a processor. As shown in the figure, the electronic device 400 includes: Figure 4
[0272] The memory 410 is configured to store a computer program executable on the processor 420.
[0273] The processor 420 is configured to execute the program stored in the memory 410 to implement the above-mentioned processing method applied to the processing device.
[0274] A computer program is provided, including computer readable codes, when the computer readable codes are executed in a computer device, a processor in the computer device performs part or all steps of the above-mentioned processing method.
[0275] A computer program product is provided, including a computer program or instructions, when the computer program or instructions are executed by a processor, part or all steps of the above-mentioned processing method are implemented.
[0276] A computer readable storage medium is provided, which stores a computer program, and the computer program can be executed by a processor to implement the above-mentioned processing method.
[0277] The present application is described with reference to the flowcharts and / or block diagrams of the method, device and electronic device according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general purpose computer, a special purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the computer or other programmable data processing device produce a device for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The device for implementing the functions specified in one or more flows and / or blocks Figure 1 The device for implementing the functions specified in one or more flows and / or blocks
[0278] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0279] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions that are executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow or flows and / or blocks Figure 1 function specified in the flow or flows and / or blocks Figure 1 of the block or blocks.
[0280] The above description is merely that of a specific implementation of the application, and the protection scope of the application is not limited thereto. Any changes or substitutions easily conceived by those skilled in the art within the technical scope disclosed by the application should be covered within the protection scope of the application.
[0281] The above description of the device embodiments is similar to the above description of the method embodiments, and has similar beneficial effects to the method embodiments. For technical details not disclosed in the device embodiments of the application, please refer to the description of the method embodiments of the application.
[0282] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily mean the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that the size of the serial number of each process in various embodiments of the application does not mean the execution order, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application. The serial number of the above embodiments of the application is only for description, not representing the advantages and disadvantages of the embodiments.
[0283] It should be noted that, in the present document, the terms "comprising", "containing", or any other similar term are intended to encompass non-exclusive inclusion, such that processes, methods, articles, or apparatuses that comprise a list of elements are not limited to those elements, but can also include other elements not expressly listed, or also include elements inherent in such processes, methods, articles, or apparatuses. Without further limitation, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0284] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed components can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0285] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units; they can be located in one place or distributed on multiple network units; and part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0286] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be realized in the form of hardware or in the form of hardware plus software functional unit.
[0287] Those of ordinary skill in the art can understand that all or part of the steps of the above method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program executes the steps of the above method embodiments when executed; and the foregoing storage medium includes: mobile storage device, read only memory (Read Only Memory, ROM), magnetic disc or optical disc, and various storage program codes.
[0288] Alternatively, the above-mentioned integrated units of the present application, if realized in the form of software function modules and sold or used as independent products, can also be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the embodiments of the present application. The aforementioned storage medium includes: mobile storage devices, ROM, magnetic disks or optical disks, and various media that can store program codes.
[0289] The above is only an embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application.
Claims
1. A processing method comprising: In response to receiving a first operation, obtaining a target parameter corresponding to the first operation; The first operation includes an editing operation on a first delivery address, and the target parameter includes editing information of the first delivery address and / or an operation parameter of the first operation; determining a target credibility of the first operation based on the target parameter; The target credibility is used to determine whether the first operation is abnormal.
2. The method according to claim 1, wherein determining the target credibility of the first operation based on the target parameter comprises: determining a target characteristic between the first delivery address and the second delivery address based on the target parameter; The second delivery address is obtained after performing the editing operation on the first delivery address; determining a target credibility of the first operation based on the target feature; Among them, the target feature includes a first feature, a second feature, and / or a third feature, the first feature is used to characterize the degree of matching between the first delivery address and the second delivery address in the semantic dimension and the spatial dimension, the second feature is used to describe the text change process between the first delivery address and the second delivery address, and the third feature is used to characterize the operation frequency of the editing operation.
3. The method according to claim 2, wherein determining the target feature between the first delivery address and the second delivery address based on the target parameter comprises: determining a target semantic difference between the first delivery address and the second delivery address based on the target parameter; determining a target spatial similarity between the first delivery address and the second delivery address based on the target parameter; The first feature is determined based on the target semantic difference and the target spatial similarity.
4. The method according to claim 3, wherein the editing information of the first delivery address includes editing process information, and the editing process information includes an editing distance between the first delivery address and the second delivery address; The determining, based on the target parameter, a target semantic difference between the first delivery address and the second delivery address includes: determining an edit distance between the first delivery address and the second delivery address; Determining vector similarity between a first text vector and a second text vector, wherein the first text vector is a text vector corresponding to the first delivery address, and the second text vector is a text vector corresponding to the second delivery address; The target semantic difference is determined based on the edit distance and the vector similarity.
5. The method according to claim 3, wherein the edit information of the first delivery address includes edit result information, and the edit result information includes the second delivery address obtained by editing the first delivery address; The determining, based on the target parameter, the target space similarity between the first delivery address and the second delivery address includes: Determining a spherical distance between the first delivery address and the second delivery address based on the geographic location information corresponding to the first delivery address and the second delivery address respectively; The target space similarity is determined based on the spherical distance and a preset distance baseline value; the distance baseline value is used to determine the degree of similarity between the first delivery address and the second delivery address in geographical space.
6. The method according to claim 2, wherein determining the target feature between the first delivery address and the second delivery address based on the target parameter comprises: determining, based on the target parameter, a character change feature between a first text corresponding to the first delivery address and a second text corresponding to the second delivery address; determining, based on the target parameter, a feature of an administrative region change between a first text corresponding to the first delivery address and a second text corresponding to the second delivery address; The second feature is determined based on the character change feature and the administrative area change feature.
7. The method according to claim 6, wherein the editing operation comprises an editing action sequence, the editing action comprising an add action, a delete action, or a modify action, the operation parameter of the first operation comprises a target ratio of the modify action in the editing action sequence, and the editing information of the first delivery address comprises editing process information, the editing process information comprising the number of modified target characters in the first text and the position of the target characters in the first text; The determining, based on the target parameter, a character change feature between a first text corresponding to the first delivery address and a second text corresponding to the second delivery address includes: determining a modification concentration of the first text based on a position identification value of each target character in the first text, an average of the position identification values of all the target characters, and the number of the target characters; The character change feature is determined based on the target proportion and the modification concentration.
8. The method according to claim 6, wherein the edit information of the first delivery address includes edit result information, and the edit result information includes administrative area information of the second delivery address obtained after editing the first delivery address; The determining, based on the target parameter, a feature of an administrative region change between a first text corresponding to the first delivery address and a second text corresponding to the second delivery address includes: Determining, from the second text, second administrative region information that is different from first administrative region information of the same administrative region level in the first text; determining the administrative region change feature based on the administrative region transfer probabilities corresponding to the administrative region information of each administrative region level in the second administrative region information; For the administrative region information of each administrative region level, the corresponding administrative region transfer probability is the probability of transferring from the administrative region information of the upper level of the administrative region level to the administrative region information of the administrative region level.
9. The method according to claim 8, further comprising: Obtaining a target address data set; the target address data set includes a target number of target delivery addresses; For each first administrative region level of administrative region information, determine a ratio of a first number of corresponding first-level addresses to a second number of target upper-level addresses as the administrative region transfer probability of the administrative region information of the first administrative region level; The first-level address is the target delivery address that contains the administrative area information of the first administrative area level in the target address data set, and the target upper-level address is the target delivery address that contains the administrative area information of the upper administrative area level corresponding to the administrative area information of the first administrative area level.
10. The method according to claim 9, further comprising: For each administrative region information of the second administrative region level, determine the ratio of the third number of corresponding second-level addresses to the target number as the administrative region transfer probability of the administrative region information of the second administrative region level; The second-level address is the target delivery address in the target address data set that contains administrative area information of the second administrative area level; the second administrative area level is the highest administrative area level, and the second administrative area level is higher than the first administrative area level.