An information mark generation method, system, device and computer storage medium

CN122838403APending Publication Date: 2026-09-29FOUNDER SECURITIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611275679.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0003]当前,各第三方资讯供应商均自行维护一套资讯标识(ID)生成规则,其格式、长度、编码逻辑各不相同,且无统一的来源标识;自建资讯也缺乏标准化的ID生成规则,导致内部系统无法直接使用供应商提供的ID作为全局唯一标识

Benefits of technology

[0016]本申请提供的一种资讯标识生成方法,获取目标资讯的发布时间和来源信息;根据来源信息,生成目标资讯的供应商资讯编码;对发布时间进行时间数值转换,得到时间数值;获取当前随机数位数;根据当前随机数位数,对时间数值和生成的随机数进行位拼接运算,生成核心标识串;提取发布时间中的年份信息,将年份信息与核心标识串拼接,得到基础标识;生成基础标识的校验位,并将校验位附加于基础标识的末尾,得到目标资讯的中台标识;建立供应商资讯编码与中台标识间的一一对应关系并存储。本申请中,通过供应商资讯编码实现了多来源资讯的标准化标识和来源区分,消除了跨供应商标识冲突的可能性;通过动态获取随机数位数,使标识的位宽分配能够随系统并发压力自适应变化;通过位拼接运算在固定总长度内生成核心标识串,确保了标识在前端系统中的数值精度兼容性;通过将年份信息与核心标识串拼接,使中台标识携带年份语义,实现了按年分库分表和高效归档;通过校验位机制实现了标识完整性的快速自检,无需查询数据库即可识别数据损坏;通过供应商资讯编码与中台标识的一一对应关系实现了资讯跨系统的双向追溯和对账,便于对资讯进行高效、准确的管理。本申请提供的一种资讯标识生成系统、电子设备及计算机可读存储介质也解决了相应技术问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838403A_ABST
    Figure CN122838403A_ABST
Patent Text Reader

Abstract

The application discloses a kind of information identification generation method, system, equipment and computer storage medium, it is related to communication and data processing technical field, obtains the publishing time and source information of target information;According to source information, the supplier information code of target information is generated;Time numerical conversion is carried out to publishing time, and time numerical value is obtained;Current random number digit is obtained;According to current random number digit, bit splicing operation is carried out to time numerical value and generated random number, and core identification string is generated;Year information in publishing time is extracted, and year information is spliced with core identification string, to obtain basic identification;The check digit of basic identification is generated, and check digit is attached to the end of basic identification, to obtain the mid-platform identification of target information;One-to-one correspondence between supplier information code and mid-platform identification is established and stored.The application eliminates the possibility of cross-supplier identification conflict, realizes per year database and efficient archiving, and facilitates efficient and accurate management of information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication and data processing technology, and more specifically, to an information identifier generation method, system, device, and computer storage medium. Background Technology

[0002] With the development of information services, the information platform needs to connect to multiple third-party information providers, while also supporting self-built information by the operating platform, so as to achieve unified entry, management and distribution of all information.

[0003] Currently, each third-party information provider maintains its own set of information identifier (ID) generation rules, which vary in format, length, and encoding logic, and lack a unified source identifier. Self-built information also lacks standardized ID generation rules, making it impossible for internal systems to directly use supplier-provided IDs as globally unique identifiers. Directly using the supplier's original ID presents problems such as inconsistent rules and the potential for ID duplication across different suppliers. Using database auto-incrementing IDs or UUIDs (Universally Unique Identifiers) as internal identifiers presents challenges: auto-incrementing IDs lack business semantics, are difficult to partition in databases, and UUIDs are too long, unordered, and have poor indexing performance. While using the Snowflake algorithm to generate internal IDs achieves global uniqueness and trend-incrementing, it cannot be compatible with standardized associations of supplier IDs, does not distinguish information sources, and the IDs do not carry year information.

[0004] In conclusion, how to manage information efficiently and accurately is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide an information identifier generation method, which can, to a certain extent, solve the technical problem of how to manage information efficiently and accurately. This application also provides an information identifier generation system, an electronic device, and a computer-readable storage medium.

[0006] To achieve the above objectives, this application provides the following technical solution: Firstly, a method for generating information identifiers is provided, including: Obtain the release time and source information of the target information; Based on the source information, generate the supplier information code for the target information; The publication time is converted into a time value to obtain a time value; Get the number of digits in the current random number; Based on the current number of bits in the random number, perform a bit concatenation operation on the time value and the generated random number to generate a core identifier string; Extract the year information from the publication time, and concatenate the year information with the core identifier string to obtain the basic identifier; Generate a check bit for the basic identifier and append the check bit to the end of the basic identifier to obtain the middleware identifier for the target information; Establish and store a one-to-one correspondence between the supplier information code and the middle platform identifier.

[0007] On the other hand, after appending the verification bit to the end of the basic identifier to obtain the middleware identifier of the target information, the method further includes: Obtain a Bloom filter; The Bloom filter is used to detect whether there are conflicting identifiers of the middle platform identifier; If a conflict exists between the middleware identifiers, return to the step of obtaining the current number of random bits; If there is no conflicting identifier with the middle platform identifier, then the one-to-one correspondence between the supplier information code and the middle platform identifier is established and stored.

[0008] On the other hand, the Bloom filter includes a first-stage Bloom filter and a second-stage Bloom filter; the first-stage Bloom filter is a count Bloom filter, used to store information identifiers generated within the most recent first preset time period; the second-stage Bloom filter is a standard Bloom filter, used to store all historical information identifiers. The Bloom filter is used to detect whether there are conflicting identifiers in the middle platform identifier, including: Check if the middle platform identifier exists in the first-level Bloom filter; If the middle platform identifier is not present in the first-stage Bloom filter, then it is determined that there is no conflict identifier for the middle platform identifier. If the middle platform identifier exists in the first-level Bloom filter, then check if the middle platform identifier exists in the second-level Bloom filter; If the middle platform identifier is not present in the second-stage Bloom filter, then it is determined that there is no conflict identifier for the middle platform identifier. If the middle platform identifier exists in the second-stage Bloom filter, then a conflict identifier with the middle platform identifier is determined to exist.

[0009] On the other hand, obtaining the current number of bits in the random number includes: Obtain the concurrent number of identifiers and the number of bits in the monitoring random number within a preset monitoring time window; A real-time conflict rate is generated based on the concurrent number of the identifier and the number of bits in the monitoring random number. The real-time conflict rate is positively correlated with the concurrent number of the identifier and negatively correlated with the number of bits in the monitoring random number. The number of digits in the current random number is determined based on the real-time conflict rate.

[0010] On the other hand, determining the current number of bits in the random number based on the real-time collision rate includes: Get the number of bits in the historical random number from the previous moment; If the real-time conflict rate within a preset time period is greater than a first threshold, and the number of bits in the historical random number is less than the maximum number of bits threshold, then the number of bits in the historical random number is increased by a preset step size to obtain the current number of bits in the random number. If the real-time conflict rate within the preset time period is less than the second threshold, and the number of bits in the historical random number is greater than the minimum number of bits threshold, then the number of bits in the historical random number is reduced by the preset step size to obtain the current number of bits in the random number. Wherein, the first threshold is greater than the second threshold.

[0011] On the other hand, generating the check bit of the basic identifier includes: Obtain a preset weight sequence, wherein the weight values ​​in the preset weight sequence are arranged cyclically; For each digit in the basic identifier, obtain the product of that digit and the weight value at the corresponding position in the preset weight sequence; If the product value is greater than or equal to the first value, then the second value is subtracted from the product value to obtain the adjustment value; Sum all the adjusted values ​​to obtain a weighted sum; Based on the first value, the weighted sum is moduloed to obtain the modulo result; The check bit is calculated based on the modulo result, and the sum of the weighted sum and the check bit is an integer multiple of the first value.

[0012] On the other hand, based on the source information, a supplier information code for the target information is generated, including: Obtain the supplier type identifier from the source information; If the supplier type identifier is a third-party supplier, then the fixed prefix corresponding to the supplier type identifier is queried from the preset prefix mapping table, and the fixed prefix is ​​concatenated with the third-party supplier's own information identifier through a separator to obtain the supplier information code of the target information; If the supplier type is identified as a self-built platform, the preset self-built prefix is ​​concatenated with the millisecond value of the current system time to obtain the supplier information code of the target information.

[0013] Secondly, an information identifier generation system is provided, including: The information acquisition module is used to obtain the publication time and source information of the target information; The supplier code generation module is used to generate a supplier information code for the target information based on the source information. The time conversion module is used to convert the release time into a time value; The bit depth acquisition module is used to obtain the number of bits in the current random number. The identifier string generation module is used to perform bit concatenation operations on the time value and the generated random number based on the current number of bits in the random number to generate the core identifier string; The splicing module is used to extract the year information from the release time, and splice the year information with the core identifier string to obtain the basic identifier; The middle platform identifier generation module is used to generate the check bit of the basic identifier and append the check bit to the end of the basic identifier to obtain the middle platform identifier of the target information. The association module is used to establish and store a one-to-one correspondence between the supplier information code and the middle platform identifier.

[0014] Thirdly, an electronic device is provided, comprising: Memory, used to store computer programs; A processor, configured to implement the steps of the information identifier generation method as described above when executing the computer program.

[0015] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the steps of any of the above-described information identifier generation methods.

[0016] This application provides a method for generating information identifiers, which includes: obtaining the publication time and source information of target information; generating a supplier information code for the target information based on the source information; converting the publication time into a time value; obtaining the number of digits in the current random number; performing a bit concatenation operation on the time value and the generated random number based on the number of digits in the current random number to generate a core identifier string; extracting the year information from the publication time and concatenating the year information with the core identifier string to obtain a basic identifier; generating a check digit for the basic identifier and appending the check digit to the end of the basic identifier to obtain a middleware identifier for the target information; and establishing and storing a one-to-one correspondence between the supplier information code and the middleware identifier. This application achieves standardized identification and source differentiation of multi-source information through supplier information coding, eliminating the possibility of cross-supplier identifier conflicts; dynamically obtaining random numbers allows the identifier's bit width allocation to adapt to system concurrency pressure; generating a core identifier string within a fixed total length through bit concatenation operations ensures numerical precision compatibility of the identifier in the front-end system; concatenating year information with the core identifier string enables the middle platform identifier to carry year semantics, achieving year-based database and table partitioning and efficient archiving; a check bit mechanism enables rapid self-checking of identifier integrity, identifying data corruption without querying the database; and the one-to-one correspondence between supplier information codes and middle platform identifiers enables bidirectional traceability and reconciliation of information across systems, facilitating efficient and accurate information management. The information identifier generation system, electronic device, and computer-readable storage medium provided in this application also solve corresponding technical problems. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0018] Figure 1 A flowchart illustrating an information identifier generation method provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of an information identifier generation system provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 4 This is another structural schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] Please see Figure 1 , Figure 1 This is a flowchart of an information identifier generation method provided in an embodiment of this application.

[0021] This application provides an information identifier generation method that can be applied to an information middleware server to achieve standardized generation, unified management, and efficient conflict detection of information identifiers in multi-vendor information access scenarios. The method may include the following steps: Step S101: Obtain the release time and source information of the target information.

[0022] In practical applications, when a piece of information is collected or created and is ready to be stored in the information platform, the first step is to obtain the publication time and source information of the target information. The publication time refers to the actual time when the information is published, such as the specific date and time of a news article, which can be represented in a standard date and time format. The source information is used to identify the source channel of the information, including but not limited to the name of a third-party supplier or the identifier of a self-built platform.

[0023] In real-world business scenarios, the publication time of news may be passed in various formats. For news pushed by third-party vendors, the publication time is usually included in the news data package provided by the vendor; for news built by the operations platform itself, the publication time can be the time entered by the operations staff. If the operations staff does not enter the time, the current time is used as the default publication time. Furthermore, for consistency in subsequent processing, after obtaining the publication time, it can be uniformly converted into a standard time object, such as a Java standard event object, to facilitate subsequent extraction of year information and time value conversion.

[0024] Understandably, the release time is one of the core input parameters for the subsequent generation of the middle platform identifier. This is because the middle platform identifier needs to carry year information to support year-based database sharding and table partitioning. At the same time, the core identifier string of the middle platform identifier needs to be encoded based on the value of the release time to ensure the increasing trend of the identifier. That is, the later the release time, the larger the value of the generated middle platform identifier. This is beneficial for the sequential writing of database indexes, reduces page splits, and improves write performance.

[0025] Step S102: Generate the supplier information code for the target information based on the source information.

[0026] In practical applications, based on the source information obtained, the supplier information code (infoCode) of the target information can be generated according to the preset standardized rules. The supplier information code refers to the unique information identifier at the supplier level generated when the supplier collects information or the operating platform creates self-built information. The supplier information code can include supplier prefix information to intuitively reflect the source of the information.

[0027] In an exemplary embodiment, the logic for generating supplier information codes can be divided into two scenarios: the first is third-party supplier information, and the second is self-built information. Specifically, in the process of generating supplier information codes, the supplier type identifier in the source information can be obtained first. This supplier type identifier is an enumerated value or string used to distinguish different source channels, such as DC representing supplier A, XGB representing supplier B, SELF representing self-built platform, etc.

[0028] Next, if the supplier type identifier is a third-party supplier, the system retrieves the fixed prefix corresponding to that supplier type identifier from the preset prefix mapping table. The prefix mapping table is a pre-configured data structure that records the correspondence between supplier types and fixed prefixes. Its configuration can be defined in the system configuration file or stored in the configuration center to achieve dynamic updates without restarting the service. After obtaining the fixed prefix, the system obtains the supplier's own information identifier assigned to the information. This supplier's own information identifier is a unique identifier maintained by each supplier within its own system. Its length and format are defined by each supplier and can be pure numbers, pure letters, or a mixture of numbers and letters. The format of the supplier's own information identifier is not strictly enforced; it is treated as an opaque string. The fixed prefix and the supplier's own information identifier are then concatenated using a separator to form the supplier information code for the target information. The separator can be a hyphen "-", and the format can be {supplier prefix}-{supplier's own information identifier}, etc. For example, if a piece of information from supplier A has a supplier's own information identifier of 12345678, then the corresponding supplier information code is DC-12345678. In this way, different supplier identifiers are isolated by using a fixed prefix. Even if different suppliers have the same information identifier, the generated supplier information codes will inevitably be different due to the different prefixes, fundamentally eliminating the possibility of cross-supplier identifier conflicts. At the same time, the intuitiveness of the prefix allows maintenance personnel and business systems to quickly identify the source of information without querying additional fields, improving the maintainability and observability of the system.

[0029] Finally, if the supplier type is identified as a self-built platform, an independent generation rule is used. Specifically, a preset self-built prefix is ​​obtained, which can be xf-ACT, where xf represents the self-built platform identifier and ACT represents self-built information; the current system time's millisecond-level timestamp is obtained, converted into a string, and concatenated with the self-built prefix to form the supplier information code, in the format xf-ACT{timestamp}. For example, if the current system time is X year X month X day X hour X minute X second X millisecond, its corresponding millisecond timestamp is 1713161746789, then the supplier information code is xf-ACT1713161746789. In this way, because millisecond-level timestamps are used as the core component of the information encoding for self-built information providers, if System.currentTimeMillis() is called twice consecutively within the same instance, the timestamp value returned by the second call will not be less than that of the first. Therefore, the combination of the xf-ACT prefix and the uniquely incrementing timestamp can guarantee uniqueness within the scope of self-built information. Even if multiple self-built information items are generated within the same millisecond, the probability of creating a large number of self-built information items within the same millisecond in actual scenarios is extremely low because the information creation action itself takes time. Therefore, uniqueness can be guaranteed in actual business.

[0030] Step S103: Convert the release time into a time value to obtain the time value.

[0031] In practical applications, the obtained publication time needs to be converted into a time value. For example, obtain the millisecond-level timestamp corresponding to the publication time, and then divide the millisecond-level timestamp by 1000 to convert it into a second-level timestamp. The reason for using a second-level timestamp instead of a millisecond-level timestamp is that in the identifier generation algorithm, a second-level timestamp is sufficient to ensure the increasing trend of the identifier, while reserving more bit width space for the random number part, which helps to reduce the probability of collision. In this way, through time value conversion, the time information can be mapped to an integer. This integer is then combined with the random number through bitwise operations to generate a unique identifier string that contains both time information and random components. Bitwise operations are one of the most efficient operations in computer systems. They are all completed in the CPU registers and do not involve memory access or I / O operations. Therefore, they have extremely high generation efficiency and can support tens of thousands or even hundreds of thousands of identifier generation requests per second.

[0032] Step S104: Obtain the current number of digits in the random number.

[0033] In practical applications, it is necessary to obtain the current random number bit length currentRandomBit. The current random number bit length is not fixed, but is dynamically maintained based on the real-time monitoring of the collision rate feedback, and its value range is 15 to 28 bits.

[0034] In an exemplary embodiment, the generation of the core identifier string in the middleware identifier relies on the number of random numbers (RANDOM_BIT). The value of the random number bit directly determines two key indicators: first, the size of the identifier's random space, which directly affects the collision probability; and second, the ratio of the timestamp to the random number bit width in the identifier, which indirectly affects indexing efficiency. If RANDOM_BIT is fixed at a large value, although the collision probability is extremely low in high-concurrency scenarios, the random number portion occupies too much bit width in low-concurrency scenarios, resulting in wasted identifier space. Conversely, if RANDOM_BIT is fixed at a small value, it performs well in low-concurrency scenarios, but once high-concurrency traffic is encountered, the collision probability will rise sharply, leading to a large number of collision retries. To balance these two indicators, this application designs an adaptive random number bit width dynamic adjustment mechanism, that is, dynamically adjusting the random number bit width according to the monitored real-time identifier collision rate. In low-concurrency scenarios, the bit width is reduced to save identifier space and improve indexing efficiency, while in high-concurrency scenarios, the bit width is increased to reduce the collision rate, achieving adaptive matching between system resources and concurrency pressure.

[0035] Based on this, during the process of obtaining the current number of random bits, the concurrent flags and the number of monitoring random bits within a preset monitoring time window can be obtained. The monitoring time window can be 60 seconds, etc. The concurrent flags refer to the total number of attempts to generate platform flags within this monitoring time window, including the number of initial generation and regeneration. This value reflects the real-time concurrency pressure of the system. The number of monitoring random bits refers to the number of random bits actually used within this time window, i.e., the RANDOM_BIT value at the current moment. Based on the concurrent flags and the number of monitoring random bits, a real-time conflict rate is generated. The real-time conflict rate is positively correlated with the concurrent flags and negatively correlated with the number of monitoring random bits. The real-time conflict rate P can be obtained through... Generate, where n is the concurrent identifier and R is the number of random numbers; obtain the historical number of random numbers from the previous moment, which is the number of random numbers used in the previous monitoring time window; if the real-time conflict rate within a preset duration is greater than a first threshold and the historical number of random numbers is less than the maximum number of random numbers threshold, then increase the historical number of random numbers by a preset step size to obtain the current number of random numbers; if the real-time conflict rate within a preset duration is less than a second threshold and the historical number of random numbers is greater than the minimum number of random numbers threshold, then decrease the historical number of random numbers by a preset step size to obtain the current number of random numbers; where the first threshold is greater than the second threshold.

[0036] For ease of understanding, let's assume a preset duration of 1 minute, a first threshold of 0.1%, a maximum random number bit threshold of 28 bits, and a preset step size of 2 bits. In the process of increasing the historical random number bit size by the preset step size to obtain the current random number bit size, if the collision rate exceeds 0.1%, it indicates that the current random space is insufficient to support the current concurrency, and the random number bit size needs to be increased to expand the random space. For every 2 bits added, the random space expands to four times its original size, and the collision probability theoretically decreases to approximately one-quarter of its original value. The step size is chosen to be 2 bits instead of 1 bit to ensure sufficient adjustment force and avoid frequent minor adjustments. Adjustment; Assuming the second threshold is 0.01% and the minimum number of bits for the random number is 15, then in the process of reducing the number of bits of the historical random number by a preset step size to obtain the current number of bits, if the real-time conflict rate is below 0.01% for 10 consecutive monitoring time windows (i.e., 10 minutes), the adjustment will be triggered. This introduces a lag effect, which can prevent frequent oscillations of the number of bits of the random number due to occasional traffic fluctuations, thus maintaining the stability of the scheme. Since every reduction of 2 bits reduces the bit width occupied by the random number part in the identifier, it is equivalent to increasing the bit width of the timestamp part, making the identifier more sequential and improving the index write performance. Ultimately, in low-concurrency scenarios, such as less than 100 TPS (Transactions Per Second), the number of bits for the random number is reduced to 15, resulting in shorter IDs and higher indexing efficiency; in medium-concurrency scenarios of 100-1000 TPS, the number of bits for the random number is maintained at 21, balancing ID length and conflict rate; and in high-concurrency scenarios of more than 1000 TPS, the number of bits for the random number is increased to 28, resulting in an extremely low conflict rate.

[0037] Step S105: Based on the number of digits in the current random number, perform a bit concatenation operation on the time value and the generated random number to generate the core identifier string.

[0038] In practical applications, during the generation of the middleware identifier (infoNo), a random number randomNum can be generated first. Then, based on the number of bits in the current random number currentRandomBit, the previously obtained time value timeSec is concatenated with this random number to generate the core identifier string infoNoCore. The middleware identifier is a globally unique identifier generated within the system when information is added to the information middleware database. The middleware identifier may include the year the information was published and ends with a check digit, used to uniquely identify a piece of information within the system.

[0039] In an exemplary embodiment, the process of bit concatenation operation can be expressed as: infoNoCore=((timeSec<<currentRandomBit)|randomNum)&INFO_NO_MAX; wherein, timeSec represents a second-level timestamp, which is a positive integer; currentRandomBit represents the number of bits of the current random number, for example, the current value is 21; randomNum represents the generated random number, which can be generated by RandomUtils.nextLong(0L, randomMax), wherein randomMax=~(-1L<<currentRandomBit), the meaning of -1L in binary is that all bits are 1. After shifting left by currentRandomBit bits, the lower currentRandomBit bits become 0, the higher bits are shifted out, and after bitwise inversion, a value with the lower currentRandomBit bits all being 1 and the higher bits all being 0 is obtained; timeSec<<currentRandomBit represents shifting the second-level timestamp left by currentRandomBit bits. The meaning of left shift operation is to multiply an integer by 2 to the power of currentRandomBit. Through left shift, currentRandomBit bits are vacated at the lower bits of the timestamp, and these vacated bits will be used to accommodate the random number. The purpose is to place the timestamp and the random number in different bit segments of the identifier respectively, and realize the conflict-free combination of the two through bit operation; | represents bitwise OR operation, that is, performing bitwise OR operation on the left-shifted timestamp and the random number. Since the lower currentRandomBit bits of the left-shifted timestamp are all 0, and the higher bits of the random number, that is, the part exceeding currentRandomBit bits, are all 0, the bitwise OR operation is equivalent to filling the random number into the lower currentRandomBit bits of the timestamp, realizing seamless concatenation of the two. The essence is to encode the timestamp and the random number into the same integer, so that the final result contains both time sequence information and random uniqueness information; INFO_NO_MAX represents the maximum allowable value of the core identifier string, which is ~(-1L<<INFO_NO_BIT), wherein INFO_NO_BIT can be fixed at 53 to ensure that the identifier can be accurately parsed in the front-end system. The calculation principle of this operation is the same as that of the above randomMax, and a mask with the lower 53 bits all being 1 and the higher bits all being 0 is obtained. Through the bitwise AND operation with INFO_NO_MAX, it is ensured that the final core identifier string does not exceed 53 bits. In this way, through the above bit concatenation operation, the generation of the core identifier string can be completed in a pure memory environment without accessing a database or remote service, thereby achieving high generation throughput of the value.

[0040] It's important to note that the reason INFO_NO_BIT is fixed at 53 bits is due to the constraint on the safe integer bit length in the IEEE 754 double-precision floating-point specification. Since the information identifier needs to be directly output to the web frontend or app for display and subsequent operations, if the identifier exceeds 53 bits, JavaScript will be unable to accurately represent the value, leading to precision loss and potentially causing serious problems such as incorrect information location and data association failures. Therefore, strictly limiting the core identifier string to within 53 bits ensures the precision integrity of the identifier throughout the entire chain of backend generation, storage, transmission, and frontend parsing.

[0041] Step S106: Extract the year information from the release time, and concatenate the year information with the core identifier string to obtain the basic identifier.

[0042] In practical applications, the year information can be extracted from the publication time and then concatenated with the core identifier string to obtain the basic identifier. The format of the basic identifier can be {year}{core identifier string}. For example, if the year is 2026 and the core identifier string is 1234567890123456, then the basic identifier is 20261234567890123456.

[0043] In an exemplary embodiment, the extraction of year information needs to take into account time zone factors, because the same moment may correspond to different years in different time zones. Therefore, the publication time can be converted to the local date using a preset time zone, and then the year can be extracted. For example, this can be achieved by using publishTime.toInstant().atZone(timeZone).toLocalDate().getYear(), where timeZone is a preset time zone object.

[0044] It's important to note that the year information is used as the prefix of the platform identifier to imbue the identifier with year semantics. When performing year-based database sharding and table partitioning, the first four digits of the identifier can directly determine the data table where the data resides, eliminating the need to query the publication time field. This improves the routing efficiency of database sharding and table partitioning. Furthermore, when archiving and statistically analyzing information, quick filtering can be performed directly based on the year prefix of the platform identifier. It should also be noted that the concatenation of the year information with the core identifier string is a string concatenation, not a numerical concatenation. This is to place the year information in decimal form at the beginning of the identifier, facilitating user readability and year-based string prefix matching queries.

[0045] Step S107: Generate the check bit of the basic identifier and append the check bit to the end of the basic identifier to obtain the middleware identifier of the target information.

[0046] In practical applications, a check digit can be calculated for the basic identifier and appended to the end of the basic identifier to form a complete middleware identifier.

[0047] In an exemplary embodiment, the calculation method for the check digit can employ the Luhn algorithm, also known as the modulo-10 algorithm. Since the traditional Luhn algorithm was originally designed for credit card number verification, while it can detect single-digit errors and adjacent-digit swapping errors, it is not suitable for generating the middleware identifier. Therefore, the traditional Luhn algorithm is adapted for this purpose. Specifically, during the generation of the middleware identifier in this application, a preset weight sequence can be obtained first. This weight sequence can be a cyclic sequence, such as [2, 1, 2, 1, 2, 1, 2, 1], etc. That is, for the first digit in the basic identifier, weight 2 is used; for the second digit, weight 1 is used; for the third digit, weight 2 is used; ..., for the ninth digit, weight 2 is used again, and so on. In this way, the weights are alternately arranged, making the weights of adjacent digits different, thereby detecting adjacent-digit swapping errors, such as miswriting 12 as 21. This is because the weight allocation changes after the swap, leading to a change in the weighted sum. Simultaneously, the weight sequence does not contain 0, ensuring that each digit contributes to the calculation of the check digit.

[0048] Then, for each digit in the basic identifier, the product of that digit and the corresponding weight value in the preset weight sequence is obtained. If the product value is greater than or equal to the first value, the second value is subtracted from the product value to obtain the adjustment value. The first and second values ​​can be flexibly determined as needed. For example, the first value can be 10, and the second value can be 9. If the product value is less than the first value, the product value itself is the adjustment value, which is equivalent to adding the units digit and the tens digit of the product. The purpose is to map the multiplication result back to the range of 0-the second value to avoid the weighted sum being too large.

[0049] Next, all adjustment values ​​are summed to obtain a weighted sum. Finally, the weighted sum is moduloed by the first value to obtain the modulo result. Then, a check digit is calculated based on the modulo result. The check digit ensures that the sum of the weighted sum and the check digit is an integer multiple of the first value. For example, the formula for calculating the check digit is checkDigit=(10-(sum%10))%10, where sum is the weighted sum. When sum%10=0, the check digit is 0; otherwise, the check digit is 10 minus sum%10. The calculated check digit is then converted into a string and appended to the end of the base identifier to obtain the complete middleware identifier. For example, if the base identifier is 20261234567890123456 and the calculated check digit is 7, then the complete middleware identifier is 202612345678901234567.

[0050] In this way, because the middleware identifier contains a check bit, when a middleware identifier is received, its integrity can be quickly verified without querying the database. For example, before using the identifier, such as when querying parameters according to the middleware identifier, consuming from the message queue, or reading from the cache, the last bit of the middleware identifier can be used as the actual check bit, and all the preceding bits can be used as the base identifier to recalculate the check bit. The two are compared to see if they are consistent. If they are inconsistent, it indicates that the middleware identifier has been corrupted during transmission or storage, and processing can be immediately rejected and an exception log can be recorded, thereby avoiding business logic errors caused by data corruption. The verification mechanism in this embodiment can complete the verification in O(n) time without external dependencies, where n is the length of the base identifier, i.e., 4 digits for the year plus a 53-bit core identifier string, for a total of 57 bits. It can detect single-bit errors, adjacent bit swapping errors, insertion / deletion errors, etc. A single-bit error refers to any change in any digit, an adjacent bit swapping error refers to any two adjacent digits being swapped, and an insertion / deletion error refers to inserting or deleting a digit in the identifier, causing all subsequent bits to be misaligned, effectively ensuring data quality.

[0051] It should be noted that the 64-bit ID generated by the Snowflake algorithm exceeds the safe integer range of JavaScript and cannot be accurately parsed on the front end. This application, however, pre-fixes INFO_NO_BIT to 53 bits, which limits the information identifier to within 53 bits, ensuring the accuracy and integrity of the information identifier during web front-end transmission and rendering.

[0052] Step S108: Establish and store the one-to-one correspondence between supplier information codes and middleware identifiers.

[0053] In practical applications, a one-to-one correspondence can be established between the generated supplier information codes and the generated middle platform identifiers, and this correspondence can be persistently stored. The essence of this correspondence is to establish an identity mapping between a piece of information on the supplier side and the middle platform side. This mapping follows a one-to-one principle: one supplier information code uniquely corresponds to one middle platform identifier, and one middle platform identifier uniquely corresponds to one supplier information code, ensuring that the identifiers of the same piece of information can be traced back between external and internal systems. In this way, when a business entity queries a piece of information using the supplier information code provided by the supplier, it can quickly find the corresponding middle platform identifier through the correlation mapping table, and then retrieve the complete information content in the main information table using the middle platform identifier. Conversely, when processing a piece of information using the middle platform identifier, it can also look up its corresponding supplier information code to confirm the source of the information, meeting the needs of compliance auditing and operational statistics, and solving the difficulties of reconciliation and traceability in multi-supplier information scenarios.

[0054] In the exemplary embodiment, when storing the one-to-one correspondence between supplier information codes and middleware identifiers, both fields (supplier information code and middleware identifier) ​​can be stored simultaneously in the main information table, and an additional independent association mapping table can be created specifically to store the correspondence between the two. This way, when cross-supplier reconciliation is required, the mapping table can be queried separately without scanning the entire main information table, improving query efficiency. Simultaneously, the mapping table can be independently configured with database sharding and table partitioning strategies, such as partitioning the mapping table by year, further enhancing management flexibility.

[0055] In specific application scenarios, the association mapping table can contain the following fields: id (auto-incrementing primary key), info_code (supplier information code), info_no (middle platform identifier), source_type (supplier type), gmt_create (creation time), and status (status). Among these, info_code and info_no are set as unique keys to enforce the correctness of the one-to-one relationship at the database level. This means that any attempt to insert a duplicate info_code or a duplicate info_no will be rejected by the database, thus ensuring data consistency. The status can be valid, logically deleted, etc.

[0056] In an exemplary embodiment, during the actual business operation of information, a piece of information goes through multiple state stages such as creation, review, publication, revision, removal / archiving, and recovery. Currently, the generated identifiers do not carry information about the current status of the information, such as whether it is valid, whether it has been revised, or whether it has been removed. This requires additional queries to the status table, increasing query overhead and system complexity. To solve this problem, during the process of concatenating the year information with the core identifier string to obtain the basic identifier, the current status code of the target information can be obtained. For example, status code 00 indicates that the information has generated an infoNo but has not yet been published; status code 01 indicates that the information is being distributed normally; status code 10 indicates that the information has been removed from the distribution channel; and status code 11 indicates that the information has been archived to cold storage. The basic identifier is obtained by concatenating the year information, the core identifier string, and the status code. In this way, changes in the status code will cause the check bit to be recalculated, ensuring the integrity of the status information. When the status of the target information changes, such as from published to delisted, only the status code needs to be changed and the middle platform identifier needs to be regenerated. This allows the middle platform identifier itself to carry the access control information of the information, enabling zero-query permission prediction and improving efficiency when handling large-scale information distribution.

[0057] This application provides a method for generating information identifiers, which includes: obtaining the publication time and source information of target information; generating a supplier information code for the target information based on the source information; converting the publication time into a time value; obtaining the number of digits in the current random number; performing a bit concatenation operation on the time value and the generated random number based on the number of digits in the current random number to generate a core identifier string; extracting the year information from the publication time and concatenating the year information with the core identifier string to obtain a basic identifier; generating a check digit for the basic identifier and appending the check digit to the end of the basic identifier to obtain a middleware identifier for the target information; and establishing and storing a one-to-one correspondence between the supplier information code and the middleware identifier. In this application, standardized identification and source differentiation of information from multiple sources are achieved through supplier information coding, eliminating the possibility of cross-supplier identifier conflicts; the bit width allocation of the identifier can be adaptively changed according to the system's concurrency pressure by dynamically obtaining the number of random numbers; a core identifier string is generated within a fixed total length through bit concatenation operations, ensuring the numerical precision compatibility of the identifier in the front-end system; by concatenating year information with the core identifier string, the middle platform identifier carries year semantics, realizing year-based database and table partitioning and efficient archiving; a check bit mechanism enables rapid self-checking of identifier integrity, identifying data corruption without querying the database; and the one-to-one correspondence between supplier information codes and middle platform identifiers enables bidirectional traceability and reconciliation of information across systems, facilitating efficient and accurate information management.

[0058] exist Figure 1 Based on the illustrated embodiment, in a high-concurrency business scenario of a multi-vendor information platform, tens of thousands or even hundreds of thousands of information entries may be added to the database daily, with hundreds to thousands of entry requests per second. Although the probability of identifier conflicts can be reduced by introducing random numbers, a small probability of conflicts still exists in extremely high-concurrency scenarios, such as when a large number of information entries flood in within the same second. Currently, conflict detection relies on the unique index constraints of the database; that is, when attempting to insert data, if the database reports a unique key conflict error, the identifier is regenerated and the attempt is retried. However, in high-concurrency scenarios, frequent unique index conflict detection will generate a large number of database query operations, consuming valuable database connection resources and becoming a performance bottleneck of the system. To solve this problem, this embodiment introduces a conflict pre-detection mechanism based on a Bloom filter, moving conflict detection from the database layer to the memory layer to address the aforementioned performance issues.

[0059] It's important to note that a Bloom filter is a space-efficient probabilistic data structure used to determine whether an element belongs to a set. Its core principle is as follows: it uses a bit array of length m and k independent hash functions. When adding an element to the set, the element is hashed using each of the k hash functions, resulting in k hash values. The corresponding bits in the bit array are then set to 1. When querying whether an element belongs to the set, the k hash values ​​are calculated again, and the corresponding bits in the bit array are checked to see if all bits are 1. If all bits are 1, the element may be in the set, with a certain probability of false positives. If any bit is 0, the element is definitely not in the set, with no missed detections and a false negative rate of 0. The Bloom filter's mechanism of preventing missed detections and only allowing false positives is suitable for the identifier conflict pre-detection scenario in this application. Even if a false positive occurs, judging a non-existent identifier as potentially existing only triggers an unnecessary identifier regeneration, without causing truly conflicting identifiers to be incorrectly considered conflict-free and added to the database, thus ensuring data accuracy.

[0060] Based on this, after appending the checksum to the end of the base identifier to obtain the middleware identifier of the target information, a Bloom filter can be obtained. The Bloom filter is used to detect if there are conflicting middleware identifiers. If a conflicting middleware identifier exists, the process returns to obtaining the current random number digits and subsequent steps. To avoid repeated execution, a preset retry threshold can be set, such as 3 times. If the number of consecutive retries exceeds this threshold, the loop stops and a degradation strategy can be adopted. For example, the Bloom filter pre-detection can be abandoned, and the identifier can be directly inserted into the database, relying on the database's unique index constraints for final confirmation. If the database insertion is successful, the identifier is added to the Bloom filter and the subsequent process continues. If the database reports a unique key conflict, an alarm log is recorded, and the identifier is regenerated and retried. If no conflicting middleware identifier exists, the process establishes and stores a one-to-one correspondence between the supplier information code and the middleware identifier. In this way, from the perspective of time complexity, the time complexity of conflict detection can be reduced from O(n) in the database to O(1) in memory. In high-concurrency scenarios, even if thousands of flags are generated per second, all conflict detection is completed in memory. The database is only responsible for the final data persistence writing, avoiding the database connection occupation and query overhead caused by conflict detection. This allows the database to use all its resources for data writing operations, which can improve data throughput.

[0061] In an exemplary embodiment, a two-stage Bloom filter structure can be adopted, that is, the Bloom filter includes a first-stage Bloom filter (BF_hot) and a second-stage Bloom filter (BF_all). The first-stage Bloom filter is a counting Bloom filter, used to store information identifiers generated within the most recent first preset time period, such as the last 7 days. The capacity can be 100 million elements, and the false positive rate can be set to 0.1%. The difference between the counting Bloom filter and the traditional Bloom filter is that each position of its bit array is no longer a bit, but a small counter. When an element is added, the counter at the corresponding position is incremented by 1, and when an element is deleted, the counter at the corresponding position is decremented by 1. This design enables BF_hot to support element deletion operations, thereby periodically eliminating expired hot data. The second-stage Bloom filter is a standard Bloom filter, used to store all historical information identifiers. The capacity can be 1 billion elements, and the false positive rate can be set to 1%.

[0062] When using a two-stage Bloom filter structure, during the process of detecting whether a middleware identifier has a conflict, the first-stage Bloom filter can be checked for the presence of the middleware identifier. For example, k hash values ​​of the middleware identifier can be calculated using the same k hash functions. The counter values ​​at the corresponding positions in the BF_hot bit array can be checked to see if all are greater than 0. If the middleware identifier does not exist in BF_hot, meaning at least one position has a counter of 0, according to the mathematical properties of Bloom filters, it can be directly determined that no conflict exists, because Bloom filters do not miss any cases; the conclusion that it does not exist is deterministic. If the middleware identifier exists in BF_hot... This means that the counter values ​​at all corresponding positions are greater than 0. Since BF_hot may have false positives, it cannot be directly used to determine that a conflict exists. In this case, it is necessary to further check whether the candidate middle platform identifier exists in the second-level Bloom filter BF_all. Therefore, if there is no middle platform identifier in the first-level Bloom filter, it is determined that there is no conflict identifier for the middle platform. If there is a middle platform identifier in the first-level Bloom filter, it is checked whether there is a middle platform identifier in the second-level Bloom filter. If there is no middle platform identifier in the second-level Bloom filter, it is determined that there is no conflict identifier for the middle platform. If there is a middle platform identifier in the second-level Bloom filter, it is determined that there is a conflict identifier for the middle platform.

[0063] In this way, due to the time-sensitive nature of information, most conflict detection occurs on information published within the most recent first preset time period. Therefore, the first-level BF_hot can cover the vast majority of query scenarios. Only when BF_hot returns a possibility of existence is it necessary to query BF_all, which has a larger query capacity and a higher false positive rate, for secondary confirmation. This ensures the correctness of the detection while maximizing the detection efficiency.

[0064] Please see Figure 2 , Figure 2 This is a schematic diagram of the structure of an information identifier generation system provided in an embodiment of this application.

[0065] This application provides an information identifier generation system, which may include: The information acquisition module 101 is used to acquire the publication time and source information of the target information; The supplier code generation module 102 is used to generate supplier information codes for target information based on source information; The time conversion module 103 is used to convert the release time into a time value; The bit acquisition module 104 is used to obtain the number of bits in the current random number; The identifier string generation module 105 is used to perform bit concatenation operations on the time value and the generated random number based on the current number of bits in the random number to generate the core identifier string; The splicing module 106 is used to extract the year information from the release time and splice the year information with the core identifier string to obtain the basic identifier; The middle platform identifier generation module 107 is used to generate the check bit of the basic identifier and append the check bit to the end of the basic identifier to obtain the middle platform identifier of the target information. The association module 108 is used to establish and store the one-to-one correspondence between supplier information codes and middle platform identifiers.

[0066] The information identifier generation system provided in this application embodiment may further include: The Bloom filter module is used to append a check bit to the end of the base identifier to obtain the middleware identifier of the target information, and then obtain the Bloom filter. The Bloom filter is used to detect whether there is a conflicting identifier of the middleware identifier. If there is a conflicting identifier of the middleware identifier, the step of obtaining the current random number digits is returned. If there is no conflicting identifier of the middleware identifier, the step of establishing a one-to-one correspondence between the supplier information code and the middleware identifier is executed and stored.

[0067] This application provides an information identifier generation system, in which the Bloom filter includes a first-level Bloom filter and a second-level Bloom filter; the first-level Bloom filter is a counting Bloom filter, used to store information identifiers generated within the most recent first preset time period; the second-level Bloom filter is a standard Bloom filter, used to store all historical information identifiers. The Bloom filter module is specifically used for: querying whether a middle platform identifier exists in the first-level Bloom filter; if no middle platform identifier exists in the first-level Bloom filter, then determining that there is no conflicting identifier with the middle platform identifier; if a middle platform identifier exists in the first-level Bloom filter, then querying whether a middle platform identifier exists in the second-level Bloom filter; if no middle platform identifier exists in the second-level Bloom filter, then determining that there is no conflicting identifier with the middle platform identifier; if a middle platform identifier exists in the second-level Bloom filter, then determining that there is a conflicting identifier with the middle platform identifier.

[0068] This application provides an information identifier generation system. The bit acquisition module is specifically used to: acquire the identifier concurrency and the number of monitoring random numbers within a preset monitoring time window; generate a real-time conflict rate based on the identifier concurrency and the number of monitoring random numbers, wherein the real-time conflict rate is positively correlated with the identifier concurrency and negatively correlated with the number of monitoring random numbers; and determine the current number of random numbers based on the real-time conflict rate.

[0069] This application provides an information identifier generation system. The bit acquisition module is specifically used to: acquire the historical random number bit length at the previous moment; if the real-time conflict rate within a preset time period is greater than a first threshold and the historical random number bit length is less than the maximum bit length threshold, then increase the historical random number bit length by a preset step size to obtain the current random number bit length; if the real-time conflict rate within the preset time period is less than a second threshold and the historical random number bit length is greater than the minimum bit length threshold, then decrease the historical random number bit length by a preset step size to obtain the current random number bit length; wherein, the first threshold is greater than the second threshold.

[0070] This application provides an information identifier generation system. The middleware identifier generation module is specifically used for: obtaining a preset weight sequence, wherein the weight values ​​in the preset weight sequence are arranged cyclically; for each digit in the basic identifier, obtaining the product of the digit and the weight value at the corresponding position in the preset weight sequence; if the product value is greater than or equal to a first value, subtracting a second value from the product value to obtain an adjustment value; accumulating all adjustment values ​​to obtain a weighted sum; taking the modulo of the weighted sum according to the first value to obtain a modulo result; calculating a check bit based on the modulo result, wherein the sum of the weighted sum and the check bit is an integer multiple of the first value.

[0071] This application provides an information identifier generation system. The supplier code generation module is specifically used to: obtain the supplier type identifier in the source information; if the supplier type identifier is a third-party supplier, query the fixed prefix corresponding to the supplier type identifier from the preset prefix mapping table, and concatenate the fixed prefix with the third-party supplier's own information identifier through a separator to obtain the supplier information code of the target information; if the supplier type identifier is a self-built platform, concatenate the preset self-built prefix with the millisecond value of the current system time to obtain the supplier information code of the target information.

[0072] This application also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of the information identifier generation method provided in the embodiments of this application. Please refer to... Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0073] An electronic device provided in this application includes a memory 201 and a processor 202. The memory 201 stores a computer program, and when the processor 202 executes the computer program, it implements the steps of the information identifier generation method described in any of the above embodiments.

[0074] Please see Figure 4 Another electronic device provided in this application embodiment may further include: an input port 203 connected to the processor 202 for transmitting commands input from the outside to the processor 202; a display unit 204 connected to the processor 202 for displaying the processing results of the processor 202 to the outside; and a communication module 205 connected to the processor 202 for enabling communication between the electronic device and the outside. The display unit 204 may be a display panel, a laser scanning display, etc.; the communication method adopted by the communication module 205 includes, but is not limited to, Mobile High-Definition Link (MHL), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connection: Wireless Fidelity (WiFi), Bluetooth communication technology, Bluetooth Low Energy communication technology, and communication technology based on IEEE 802.11s.

[0075] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the information identifier generation method described in any of the above embodiments.

[0076] The computer-readable storage media involved in this application include random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs (compact disc read-only memory), or any other form of storage media known in the art.

[0077] This application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the information identifier generation method described in any of the above embodiments.

[0078] For descriptions of relevant parts of the information identifier generation system, electronic device, and computer-readable storage medium provided in the embodiments of this application, please refer to the detailed descriptions of the corresponding parts in the information identifier generation method provided in the embodiments of this application, which will not be repeated here. Furthermore, parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.

[0079] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0080] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating information identifiers, characterized in that, include: Obtain the release time and source information of the target information; Based on the source information, generate the supplier information code for the target information; The publication time is converted into a time value to obtain a time value; Get the number of digits in the current random number; Based on the current number of bits in the random number, perform a bit concatenation operation on the time value and the generated random number to generate a core identifier string; Extract the year information from the publication time, and concatenate the year information with the core identifier string to obtain the basic identifier; Generate a check bit for the basic identifier and append the check bit to the end of the basic identifier to obtain the middleware identifier for the target information; Establish and store a one-to-one correspondence between the supplier information code and the middle platform identifier.

2. The information identifier generation method according to claim 1, characterized in that, After appending the verification bit to the end of the basic identifier to obtain the middleware identifier of the target information, the method further includes: Obtain a Bloom filter; The Bloom filter is used to detect whether there are conflicting identifiers of the middle platform identifier; If a conflict exists between the middleware identifiers, return to the step of obtaining the current number of random bits; If there is no conflicting identifier with the middle platform identifier, then the one-to-one correspondence between the supplier information code and the middle platform identifier is established and stored.

3. The information identifier generation method according to claim 2, characterized in that, The Bloom filter includes a first-level Bloom filter and a second-level Bloom filter; the first-level Bloom filter is a count Bloom filter, used to store information identifiers generated within the most recent first preset time period; the second-level Bloom filter is a standard Bloom filter, used to store all historical information identifiers. The Bloom filter is used to detect whether there are conflicting identifiers in the middle platform identifier, including: Check if the middle platform identifier exists in the first-level Bloom filter; If the middle platform identifier is not present in the first-stage Bloom filter, then it is determined that there is no conflict identifier for the middle platform identifier. If the middle platform identifier exists in the first-level Bloom filter, then check if the middle platform identifier exists in the second-level Bloom filter; If the middle platform identifier is not present in the second-stage Bloom filter, then it is determined that there is no conflict identifier for the middle platform identifier. If the middle platform identifier exists in the second-stage Bloom filter, then a conflict identifier with the middle platform identifier is determined to exist.

4. The information identifier generation method according to claim 2, characterized in that, Get the number of digits in the current random number, including: Obtain the concurrent number of identifiers and the number of bits in the monitoring random number within a preset monitoring time window; A real-time conflict rate is generated based on the concurrent number of the identifier and the number of bits in the monitoring random number. The real-time conflict rate is positively correlated with the concurrent number of the identifier and negatively correlated with the number of bits in the monitoring random number. The number of digits in the current random number is determined based on the real-time conflict rate.

5. The information identifier generation method according to claim 4, characterized in that, Based on the real-time collision rate, determine the current number of bits in the random number, including: Get the number of bits in the historical random number from the previous moment; If the real-time conflict rate within a preset time period is greater than a first threshold, and the number of bits in the historical random number is less than the maximum number of bits threshold, then the number of bits in the historical random number is increased by a preset step size to obtain the current number of bits in the random number. If the real-time conflict rate within the preset time period is less than the second threshold, and the number of bits in the historical random number is greater than the minimum number of bits threshold, then the number of bits in the historical random number is reduced by the preset step size to obtain the current number of bits in the random number. Wherein, the first threshold is greater than the second threshold.

6. The information identifier generation method according to claim 1, characterized in that, Generating the checksum of the basic identifier includes: Obtain a preset weight sequence, wherein the weight values ​​in the preset weight sequence are arranged cyclically; For each digit in the basic identifier, obtain the product of that digit and the weight value at the corresponding position in the preset weight sequence; If the product value is greater than or equal to the first value, then the second value is subtracted from the product value to obtain the adjustment value; Sum all the adjusted values ​​to obtain a weighted sum; Based on the first value, the weighted sum is moduloed to obtain the modulo result; The check bit is calculated based on the modulo result, and the sum of the weighted sum and the check bit is an integer multiple of the first value.

7. The information identifier generation method according to claim 1, characterized in that, Based on the source information, a supplier information code for the target information is generated, including: Obtain the supplier type identifier from the source information; If the supplier type identifier is a third-party supplier, then the fixed prefix corresponding to the supplier type identifier is queried from the preset prefix mapping table, and the fixed prefix is ​​concatenated with the third-party supplier's own information identifier through a separator to obtain the supplier information code of the target information; If the supplier type is identified as a self-built platform, the preset self-built prefix is ​​concatenated with the millisecond value of the current system time to obtain the supplier information code of the target information.

8. An information identifier generation system, characterized in that, include: The information acquisition module is used to obtain the publication time and source information of the target information; The supplier code generation module is used to generate a supplier information code for the target information based on the source information. The time conversion module is used to convert the release time into a time value; The bit depth acquisition module is used to obtain the number of bits in the current random number. The identifier string generation module is used to perform bit concatenation operations on the time value and the generated random number based on the current number of bits in the random number to generate the core identifier string; The splicing module is used to extract the year information from the release time, and splice the year information with the core identifier string to obtain the basic identifier; The middle platform identifier generation module is used to generate the check bit of the basic identifier and append the check bit to the end of the basic identifier to obtain the middle platform identifier of the target information. The association module is used to establish and store a one-to-one correspondence between the supplier information code and the middle platform identifier.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the information identifier generation method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the information identifier generation method as described in any one of claims 1 to 7.