Data processing method and device, equipment and medium
By using the supported feature database in the service recommendation system to store the features generated by the feature production components, the problems of high access cost and low adaptability are solved, and efficient porting and general access of feature production components are achieved.
Patent Information
- Application Number
- CN202410071585.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when accessing a dedicated feature generation component to the recommendation system, component dependence and environmental adaptability need to be considered, resulting in excessive access cost and low adaptability, and it is impossible to directly access the service recommendation system of different architectures.
By using the feature database supported by the business recommendation system to store the features generated by the feature production components, simplify component dependence, reduce access costs, and realize the universality of feature production components in multiple business applications.
It reduces the access cost of feature production components, improves its versatility, reduces the secondary development requirement for feature generation components, and enhances the environmental adaptability with the business recommendation system.
Smart Images

Figure CN120335770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, and medium. Background Art
[0002] With the rapid development of Internet technology, the content on the Internet has shown an explosive growth. Users need to spend a lot of energy and time to search for the content they are interested in from the vast amount of content. In order to help users quickly obtain the information they need from the vast amount of information data, a recommendation system has emerged.
[0003] In the current recommendation scenario, it is often necessary to connect a dedicated feature generation component to the recommendation system so that the dedicated feature generation component can generate features for information recommendation, and then determine the recommended content for the user based on the features. However, when connecting the dedicated feature generation component to the recommendation system, it is usually necessary to consider the component dependencies and environmental adaptability between the dedicated feature generation component and the recommendation system. Since the architectures of recommendation systems for different business applications may be different, it is necessary to independently develop the dedicated feature generation component according to the architecture of the recommendation system to adapt to the recommendation systems of different business applications, resulting in too high access costs for the dedicated feature generation component. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, device, and medium, which can reduce the access cost of the feature generation component.
[0005] On the one hand, embodiments of this application provide a data processing method, which includes:
[0006] Obtain an initial feature set corresponding to a business application through a feature generation component, and determine a feature database corresponding to the business application according to the database parameters included in the business configuration file in the feature generation component;
[0007] Perform data format conversion on the initial object features in the initial feature set to obtain a candidate feature set; the data format of the candidate object features in the candidate feature set is the data format supported by the feature database;
[0008] Perform feature aggregation processing on the candidate object features in the candidate feature set to obtain an object aggregation feature, and store the object aggregation feature in the feature database; the feature database is used to provide a data source for the business recommendation service corresponding to the business application.
[0009] On the one hand, embodiments of this application provide a data processing apparatus, which includes:
[0010] A database determination module, configured to obtain an initial feature set corresponding to a business application through a feature production component, and determine a feature database corresponding to the business application according to database parameters included in a business configuration file in the feature production component;
[0011] A format conversion module, configured to perform data format conversion on initial object features in the initial feature set to obtain a candidate feature set; the data format corresponding to candidate object features in the candidate feature set is a data format supported by the feature database;
[0012] A feature aggregation module, configured to perform feature aggregation processing on candidate object features in the candidate feature set to obtain an object aggregation feature, and store the object aggregation feature in the feature database; the feature database is used to provide a data source for a business recommendation service corresponding to the business application.
[0013] Among them, the database determination module is specifically configured to:
[0014] Call a business data query interface to obtain an object metadata set from a business database corresponding to the business application;
[0015] Batch process object metadata in the object metadata set to obtain M data subsets; M is an integer greater than 1;
[0016] Distribute the M data subsets to M data processing nodes in the feature production component, and perform parallel processing on object metadata in the M data subsets through the M data processing nodes to obtain object output features corresponding to the M data processing nodes; object metadata in one data subset is used to be distributed to one data processing node;
[0017] Merge and process object output features corresponding to the M data processing nodes to obtain an initial feature set corresponding to the business application.
[0018] Among them, the feature aggregation module is further configured to:
[0019] If the business configuration file includes database parameters for caching the candidate feature set, determine a cache database corresponding to the feature production component; the number of candidate feature sets cached in the cache database is N, and one candidate feature set carries a generation timestamp, and N is a positive integer;
[0020] In the cache database, determine candidate feature sets whose generation timestamps belong to an aggregation period as a to-be-processed feature set;
[0021] Perform feature aggregation processing on candidate object features in the to-be-processed feature set to obtain an object aggregation feature, and clear the to-be-processed feature set in the cache database.
[0022] Among them, the feature aggregation module is specifically configured to:
[0023] Obtain the object identification information corresponding to the candidate object features in the candidate feature set, combine the candidate object features with the same object identification information to obtain the first combined feature;
[0024] Perform serialization processing on the first combined feature to obtain the first serialized feature corresponding to the first combined feature, and generate an object aggregation feature with the object identification information as the key and the first serialized feature as the value.
[0025] Among them, the feature aggregation module is specifically used for:
[0026] Obtain the key fields corresponding to the candidate object features in the candidate feature set, combine the candidate object features with the same key fields to obtain the second combined feature;
[0027] Perform serialization processing on the second combined feature to obtain the second serialized feature corresponding to the second combined feature, and generate an object aggregation feature with the key field as the key and the second serialized feature as the value.
[0028] Among them, the feature aggregation module is specifically used for:
[0029] Obtain D unit characters included in the object aggregation feature, count the occurrence frequency of each unit character in the object aggregation feature, and determine the occurrence frequency corresponding to each unit character as the weight value corresponding to each unit character; D is an integer greater than 1;
[0030] Construct a feature encoding tree corresponding to the object aggregation feature according to each unit character and the weight value corresponding to each unit character; the leaf nodes in the feature encoding tree are used to represent D unit characters; the weight value corresponding to the parent node i in the feature encoding tree is the sum of the weight values corresponding to the child nodes of the parent node i;
[0031] Obtain the shortest path between each leaf node and the root node in the feature encoding tree, and obtain the encoding information of the unit character corresponding to each leaf node according to the position information of each node in the shortest path in the feature encoding tree;
[0032] Generate an object encoding feature corresponding to the object aggregation feature according to the encoding information corresponding to each unit character and the text position of each unit character in the object aggregation feature, and store the object encoding feature in the feature database.
[0033] Among them, the data processing device further includes a database access module, and the database access module is used for:
[0034] Create a data storage class for inheriting the database access interface in the feature generation component, and create a data storage function in the data storage class;
[0035] Determine the database identification information and data storage structure corresponding to the feature database as the database parameters corresponding to the feature database;
[0036] Generate the implementation logic of the data storage function according to the database access interface and the database parameters corresponding to the feature database;
[0037] Execute the implementation logic of the data storage function and write the database parameters corresponding to the feature database into the service configuration file.
[0038] Among them, the data processing device further includes a service recommendation module, and the service recommendation module is used for:
[0039] If a service recommendation request is received, obtain the object identification information corresponding to the service recommendation request;
[0040] Determine the object aggregation features in the feature database that match the object identification information as the object recommendation features corresponding to the service recommendation request;
[0041] Obtain the set of services to be recommended corresponding to the service recommendation request, and obtain the service recommendation features corresponding to each service to be recommended in the set of services to be recommended;
[0042] Obtain the similarity evaluation values between the service recommendation features corresponding to each service to be recommended and the object recommendation features, and determine the recommendation status corresponding to each service to be recommended according to the similarity evaluation values;
[0043] Display the services to be recommended whose recommendation status belongs to the recommended enabled status in the service application.
[0044] Among them, the number of services to be recommended in the set of services to be recommended is L, and L is a positive integer; the service recommendation module is specifically used for:
[0045] Sort the similarity evaluation values corresponding to the L services to be recommended in descending order to obtain the sorted L services to be recommended;
[0046] Obtain the number y corresponding to the recommended display position in the service application. Among the sorted L services to be recommended, determine the recommendation status corresponding to the first y services to be recommended as the recommended enabled status, and determine the recommendation status corresponding to the services to be recommended other than the first y services to be recommended as the recommended disabled status; y is a positive integer less than or equal to L.
[0047] An embodiment of the present application provides a computer device on the one hand, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor executes the steps of the method in the first aspect of the embodiments of the present application.
[0048] In one aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the steps of the method in one aspect of the embodiments of the present application are executed.
[0049] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various alternative manners in the above-mentioned one aspect.
[0050] In the embodiments of the present application, the feature database supported by the business application can be used to store the object aggregation features produced by the feature production component for the business application, so as to realize transplanting and accessing the feature production component into the business recommendation system in the business application. Since the feature database is the database supported by the business application, the feature database and the business recommendation system in the business application have good environmental adaptability, which can simplify the component dependency between the feature production component and the business recommendation system, and there is no need to perform secondary development on the feature generation component, reducing the access cost of the feature production component. In addition, the feature production component can be applied to multiple business applications that support the feature database, without the need to independently develop the feature generation component according to the architecture of the business recommendation system, improving the versatility of the feature production component. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 is a schematic structural diagram of a network architecture provided by the embodiments of the present application;
[0053] Figure 2 is a schematic diagram of a business recommendation system provided by the embodiments of the present application;
[0054] Figure 3 is a schematic flowchart of a data processing method provided by the embodiments of the present application Figure 1 ;
[0055] Figure 4 is a schematic flowchart of a feature aggregation provided by the embodiments of the present application;
[0056] Figure 5 It is a schematic diagram of feature aggregation provided by an embodiment of the present application;
[0057] Figure 6 It is a schematic diagram of feature storage provided by an embodiment of the present application Figure 1 ;
[0058] Figure 7 It is a schematic flowchart of a data processing method provided by an embodiment of the present application Figure 2 ;
[0059] Figure 8 It is a schematic diagram of feature storage provided by an embodiment of the present application Figure 2 ;
[0060] Figure 9 It is a schematic diagram of feature processing provided by an embodiment of the present application;
[0061] Figure 10 It is a schematic diagram of the interface for recommending virtual skins provided by an embodiment of the present application;
[0062] Figure 11a It is a schematic diagram of feature read / write delay provided by an embodiment of the present application Figure 1 ;
[0063] Figure 11b It is a schematic diagram of feature read / write delay provided by an embodiment of the present application Figure 2 ;
[0064] Figure 11c It is a schematic diagram of feature read / write delay provided by an embodiment of the present application Figure 3 ;
[0065] Figure 12 It is a schematic diagram of the read delay of the feature database provided by an embodiment of the present application;
[0066] Figure 13 It is a schematic diagram of the business recommendation effect provided by an embodiment of the present application;
[0067] Figure 14 It is a schematic diagram of the structure of a data processing device provided by an embodiment of the present application;
[0068] Figure 15 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0070] To facilitate the understanding of the technical solutions proposed in the embodiments of the present application, the following explains the technical terms involved in the embodiments of the present application:
[0071] Flink: An open-source distributed stream computing framework written in Java (a programming language) and Scala (a programming language).
[0072] RFC (Realtime Feature Computing): A real-time feature production framework developed based on the Flink framework, which can meet the real-time data processing requirements of various businesses, covering multiple business scenarios such as real-time feature production and real-time data publishing. For example, it can be used in the business scenario of serving AMS (Ambari Metrics System, a developed cluster management system).
[0073] Remote Dictionary Server (Redis): An open-source key-value pair database written in ANSI C (a programming language).
[0074] TDBank: A real-time data acquisition system, which is a bridge between the business data source and the data processing system. It can be used to decouple the data processing system from the data source, and then provide data support for the backend distributed data warehouse (TDW, responsible for processing offline data) and the real-time computing platform.
[0075] JSON (JavaScript Object Notation): A lightweight data interchange format, which is easy for humans to read and write, and is also easy for machines to parse and generate.
[0076] FlatBuffers: An open-source serialized data library that supports multiple language interfaces.
[0077] AUC (Area Under Curve): The area under the Receiver Operating Characteristic Curve (ROC). The AUC value can be used as an evaluation criterion for classification models.
[0078] ARPU (Average Revenue Per User): The business profit contributed by each business object on average within a certain period (e.g., within one day).
[0079] Please refer to Figure 1 , Figure 1 It is a schematic structural diagram of a network architecture provided by an embodiment of this application. As Figure 1 shown, the network architecture may include a server 10d and a terminal device cluster. The terminal device cluster may include one or more terminal devices, and the number of terminal devices will not be limited here. As Figure 1 shown, the multiple terminal devices may specifically include a terminal device 10a, a terminal device 10b, and a terminal device 10c, etc.; as Figure 1 shown, the terminal device 10a, the terminal device 10b, and the terminal device 10c may be respectively connected to the server 10d through a network connection, so that each terminal device can perform data interaction with the server 10d through this network connection.
[0080] Among them, Figure 1 the shown server 10d may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The embodiment of this application does not limit the type of the server.
[0081] Figure 1 The terminal devices in the shown terminal cluster may include, but are not limited to: electronic devices such as smart phones, tablet computers, laptop computers, palmtop computers, mobile internet devices (MID), wearable devices (such as smart watches, smart bracelets, etc.), intelligent voice interaction devices, smart home appliances (such as smart TVs, etc.), in-vehicle devices, and aircraft. The embodiment of this application does not limit the type of the terminal devices.
[0082] It can be understood that, as Figure 1Any one of the terminal devices in the shown terminal cluster can install a service application. The service application can be an application client, a web page, a mini-program, etc. that can provide service recommendation services. The service application can push diverse service information to service objects. For example, information such as audio, video, news, products, virtual props, virtual characters, and virtual skins. When the service application runs on each terminal device, the server 10d can refer to the background server corresponding to the service application. When the service application is an application client, the service application running on each terminal device can be an independent client or an embedded sub-client integrated in a certain client. The embodiments of the present application do not make limitations in this regard.
[0083] In the embodiments of the present application, when the service application is an application client, the service application can specifically include but is not limited to: vehicle-mounted client, smart home client, entertainment client (such as game client, etc.), multimedia client (such as video client, music client, etc.), interactive client, and information client (such as news client, etc.). Among them, if the terminal device included in the terminal cluster is a vehicle-mounted device, then the vehicle-mounted device can be an intelligent terminal in the intelligent transportation scenario, and the application client running on the vehicle-mounted device can be called a vehicle-mounted client.
[0084] It can be understood that different service objects (users of the service application) may show different interest tendencies for different services to be recommended. To improve the accuracy of service recommendation and thus improve the experience of service objects, a service recommendation system can be integrated into the service application to preferentially display the services to be recommended that match their interest tendencies in the service application.
[0085] In the current service recommendation scenario, it is necessary to connect the feature generation framework to the service recommendation system so as to generate features for service recommendation through the feature generation framework, and then determine the recommended content for the service object based on the features. However, when connecting the feature generation component to the service recommendation system, it is usually necessary to consider the component dependencies and environment adaptability between the feature generation component and the service recommendation system. For example, the feature generation component may depend on some components of AMS, such as PageView (a component for implementing paging view effects), Marvel (a highly customized database), etc., and cannot be directly connected to the service recommendation system and needs to be redeveloped. In addition, since the architectures of the service recommendation systems of different service applications may be different, it may also be necessary to independently develop the feature generation component according to the architecture of the service recommendation system to adapt to the service recommendation systems of different service applications, resulting in too high access costs and low adaptability of the dedicated feature generation component.
[0086] To solve the above problems, in the embodiments of the present application, the feature database supported by the business recommendation system is used to store the features produced by the feature production component for business applications, so as to realize the transplantation and access of the feature production component into the business recommendation system. Since the feature database is the database supported by the business recommendation system, the feature database and the business recommendation system have good environmental adaptability, which can simplify the component dependencies between the feature production component and the business recommendation system, without the need for secondary development of the feature generation component, and reduce the access cost of the feature production component.
[0087] In addition, the embodiments of the present application propose a transplantation and access solution for the feature production component, which can transplant and access the feature production component into business applications that support multiple feature databases. For example, the feature production component can be transplanted and accessed into game applications and video applications at the same time, without the need to specifically develop a feature generation component for the game application according to the architecture of the recommendation system in the game application, and specifically develop a feature generation component for the video application according to the architecture of the recommendation system in the video application. Therefore, the access cost of the feature production component can be significantly reduced.
[0088] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a business recommendation system provided by the embodiments of the present application. As Figure 2 shown, the business recommendation system 200 may include components such as a business application server 201, a request splitting service 202, a framework service 203, a model application service 204, a feature production component 205, a feature database 206, a feature snapshot library 207, a parameter storage service (ParameterServer) 208, a model training service 209, a sample generation service 210, and a feature database filling 211, etc., and can provide online services and offline services through the above components. Among them, the feature database 206 is used to cache the object aggregation features produced by the feature production component 205, so as to realize the transplantation and access of the feature production component 205 into the business recommendation system 200.
[0089] For offline services, the business application server 201 can also record the interaction data of business objects in the business application (e.g., operation data such as browsing and transactions), so as to determine sample tags and historical interaction data (including object-side information and item-side information). The business application server 201 can send the historical interaction data to the feature database 206 through the feature library loading 211, and send the sample tags to the sample generation service 210. The sample generation service 210 obtains the backflow sample features from the feature snapshot library 207, splices and samples the sample features and sample tags, and sends the obtained sample feature set to the model training service 209. The model training service 209 learns model parameters (e.g., incremental learning, window learning, etc.) based on the sample feature set, updates the trained model parameters to the parameter storage service 208, and then publishes the model parameters to the model application service 204 through the parameter storage service 208.
[0090] For online services, the business application server 201 can interact with the terminal device installed with the business application, receive the business service request from the terminal device (which may include the object identification information corresponding to the business object, etc.), and forward the business service request to the request splitting service 202. The request splitting service 202 splits the above business service request to obtain a business sorting request, and transmits it to the framework service 203; this business sorting request can be used to indicate the sorting of multiple business services to be recommended (e.g., virtual skins, virtual props, etc.). As Figure 2 shown, the framework service 203 can include one or more plugins such as a feature plugin, a sorting plugin, and a policy plugin. These plugins can be associated with a whitelist, an AB test platform, etc. The framework service 203 can split and cache the business sorting request through the whitelist and the AB test platform, and encapsulate it into a recommendation service request that can be received and processed by the model application service 204.
[0091] The model application service 204 can pull the latest model parameters from the parameter storage service 208 and update the model parameters in the business recommendation model in the model application service 204 with the latest model parameters. The model application service 204 can obtain the object identification information corresponding to the recommendation service request (for example, account identification information), obtain the object aggregation features matching the object identification information from the feature database 206, and determine the object aggregation features as the object recommendation features corresponding to the business recommendation request; the model application service 204 can obtain the set of to-be-recommended services corresponding to the business recommendation request and obtain the business recommendation features corresponding to each to-be-recommended service in the set of to-be-recommended services; then, the object recommendation features and the business recommendation features are input into the business recommendation model, and the business recommendation model performs processing such as feature splicing, feature crossing, and feature output on the object recommendation features and the business recommendation features to score and sort the to-be-recommended services in the set of to-be-recommended services, and generate a recommended service sorted list according to the recommendation scores corresponding to the to-be-recommended services in the set of to-be-recommended services.
[0092] After obtaining the recommended service sorted list, the model application service 204 can return the recommended service sorted list to the business application server 201 through the framework service 203 and the request splitting service 202 in sequence, and the business application server 201 sends it to the terminal device so that the terminal device provides a business recommendation service for the business object according to the recommended service sorted list. Among them, the above object recommendation features and business recommendation features can also be transmitted to the feature snapshot library 207 as sample features to provide a data source for the sample generation service 210.
[0093] The data processing method involved in the embodiments of the present application will be described in detail below. Specifically, please refer to Figure 3 , Figure 3 is a flowchart of a data processing method provided by an embodiment of the present application Figure 1 , and this data processing method can be executed by a computer device, and this computer device can be Figure 1 the terminal device 10a shown in Figure 1 , or it can also be Figure 3 the server 10d shown in
[0094] Step S101: Obtain an initial feature set corresponding to a business application through a feature production component, and determine a feature database corresponding to the business application according to the database parameters included in the business configuration file in the feature production component.
[0095] Among them, the feature production component refers to a component that provides features for the business recommendation service corresponding to a business application. Specifically, it can be an RFC component, a Flink component, a Storm component, a Spark Streaming component, a Hadoop MapReduce component, a Hive component, etc., or it can be a variant of any of the above components. The feature database is the database supported by the business application, and this feature database is used to provide a data source for the business recommendation service corresponding to the business application. It can be a database that has been successfully connected to the business recommendation system in the business application. Specifically, the feature database can be a Redis database, a NoSQL database (a key-value pair database), a Voldemort database (a key-value pair database), a ClickHouse database (a columnar database), a MySQL database (a relational database), etc.
[0096] In an embodiment of the present application, the feature database supported by the business application can be connected to the feature production component to implement the transplantation and connection of the feature production component into the business recommendation system in the business application. The process of connecting the feature database to the feature production component may include: creating a data storage class for inheriting the database access interface in the feature generation component, and creating a data storage function in the data storage class; determining the database identification information and data storage structure corresponding to the feature database as the database parameters corresponding to the feature database; and then, according to the database access interface and the database parameters corresponding to the feature database, generating the implementation logic of the data storage function; executing the implementation logic of the data storage function to write the database parameters corresponding to the feature database into the business configuration file.
[0097] Among them, the database access interface (SinkFunction interface) is an abstract interface that defines the feature storage location of the feature production component. By implementing the database access interface, the feature storage location of the feature production component can be customized. The data storage class refers to a class that can inherit or implement the database access interface; the data storage function refers to a method created in the data storage class, which can be specifically used to define the storage location and storage method of the features output by the feature production component. The database parameters corresponding to the feature database may include, but are not limited to, the database identification information corresponding to the feature database, the database port information, the data storage structure (for example, the data writing frequency, the data storage type, the supported data format), etc.
[0098] Specifically, a data storage class can be created in the feature generation component, and this data storage class can inherit or implement a database access interface; a data storage function is created in the created data storage class, and the database parameters corresponding to the feature database are passed into the database access interface to generate the implementation logic of the data storage function. Further, the implementation logic of the data storage function can be executed to connect the feature database to the feature production component and write the database parameters corresponding to the feature database into the service configuration file, so as to connect the feature database to the feature production component, and further connect the feature production component to the service application.
[0099] It can be understood that when using the feature production component to generate features for a service application, the feature database corresponding to the service application can be determined according to the database parameters included in the service configuration file in the feature production component, so as to store the features output by the feature production component for the service application (for example, candidate object features and object aggregation features mentioned below) into the feature database.
[0100] The process of using the feature production component to generate an initial feature set for a service application may include: calling a service data query interface to obtain a set of object metadata from the service database corresponding to the service application; batch-processing the object metadata in the set of object metadata to obtain M data subsets; M represents the number of data subsets, M is an integer greater than 1, and the value of M can be 2, 3, 10, etc., and the specific value can be determined according to the number of data processing nodes (publish_node) included in the feature production component. Further, the M data subsets can be distributed to M data processing nodes in the feature production component, and the object metadata in the M data subsets are processed in parallel by the M data processing nodes to obtain object output features corresponding to the M data processing nodes; among them, the object metadata in one data subset is used to be distributed to one data processing node; the object output features corresponding to the M data processing nodes are merged to obtain the initial feature set corresponding to the service application. It can be seen that by processing the object metadata in parallel by multiple data processing nodes, the data processing efficiency can be significantly improved.
[0101] Among them, the initial feature set corresponding to the business application includes multiple initial object features, which are features obtained by the feature generation component after processing the object metadata corresponding to the business application. Such features can be real-time features or offline features (non-real-time features), and the specific feature type can be determined according to the actual business requirements of the business application. Object metadata refers to the interaction data generated by business objects in the business application. For example, the browsing and transaction data of business objects for historical recommendation services, the data generated during battles and ranking matches in the business application, etc. The object metadata is stored in the business database corresponding to the business application (such as Kafka database, Elasticsearch database, etc.), and this business database is the data source for the feature production component to provide feature production services. The business data query interface is the communication bridge between the feature production component and the business database. When the feature production component calls the business data query interface, it allows the feature production component to obtain object metadata from the business database through the query statements supported by the business data source.
[0102] In the embodiment of the present application, the feature production component may include multiple data processing nodes, which can be used to process object metadata. That is to say, when the feature production component executes the feature output task corresponding to the business application (such as obtaining the initial feature set corresponding to the business application), multiple data processing nodes can be specified to execute together. Specifically, the object metadata in the object metadata set can be processed in batches to obtain M data subsets. For example, the object metadata can be processed in batches according to the object identification information corresponding to the object metadata, and the object metadata with the same object identification information in the object metadata set is divided into the same data subset; or, it can also be processed in batches according to the business type corresponding to the object metadata, and the object metadata of the same business type in the object metadata set is divided into the same data subset. For example, the object metadata in data subset a is browsing-type object metadata; the object metadata in data subset b is transaction-type object metadata. It can be understood that the embodiment of the present application does not limit the data volume and data type included in each data subset, and the specific data volume and data type can be determined according to the actual business requirements.
[0103] Further, the M data subsets can be distributed to the M data processing nodes in the feature production component, and the object metadata in the M data subsets can be processed in parallel by the M data processing nodes to obtain the object output features corresponding to the M data processing nodes. Among them, the object metadata in one data subset is used to be distributed to one data processing node, and the distribution method can be random distribution or distribution according to the matching degree between the data volume of the data subset and the data processing node. For example, the data subset with a large data volume can be distributed to the data processing node with strong computing power, and the data subset with a small data volume can be distributed to the data processing node with general computing power. The embodiments of the present application do not limit the specific distribution method.
[0104] For example, the M data subsets include data subset a and data subset b. Data subset a can be distributed to data processing node A in the feature production component. Data processing node A processes data subset a according to the corresponding service processing logic, and determines the object output feature corresponding to data processing node A as the service processing result of data subset a. Data subset b can be distributed to data processing node B in the feature production component. Data processing node B processes data subset b according to the corresponding service processing logic, and determines the object output feature corresponding to data processing node B as the service processing result of data subset b. Among them, the service processing logic can be understood as the data processing basis for the object metadata in the data subset according to the service requirements corresponding to the service application. For example, the service processing logic can be "count the number of times that service object A uses virtual role B during 18:00-20:00".
[0105] After obtaining the object output features corresponding to each data processing node, the object output features corresponding to each data processing node can be merged to obtain an initial feature set corresponding to the service application. The merging method can be to perform feature splicing on the object output features corresponding to each data processing node, or it can also be to perform feature fusion on the object output features corresponding to each data processing node. For example, if the object output feature a is the object output feature corresponding to the data processing node A, and the object output feature b is the object output feature corresponding to the data processing node B, the initial feature set obtained by merging through feature splicing can be {object output feature a, object output feature b}; the initial feature set obtained by merging through feature fusion can be {object output feature c}, where the object output feature c can be the object output feature after feature fusion of the object output feature a and the object output feature b; for example, if the object output feature a is "skin_buy_cnt_1h:2 (indicating that the number of skin transactions within 1 hour is 2)", and the object output feature b is "skin_buy_cnt_1h:1 (indicating that the number of skin transactions within 1 hour is 1)", then the object output feature c obtained by feature fusion can be expressed as "skin_buy_cnt_1h:3 (indicating that the number of skin transactions within 1 hour is 3)".
[0106] Step S102: Convert the data format of the initial object features in the initial feature set to obtain a candidate feature set.
[0107] In the embodiment of the present application, the data format supported by the feature database can be obtained from the database parameters included in the service configuration file. If the data format corresponding to the initial object features in the initial feature set is not the data format supported by the feature database, the data format of the initial object features in the initial feature set needs to be converted to obtain a candidate feature set; among them, the data format corresponding to the candidate object features in the candidate feature set is the data format supported by the feature database.
[0108] For example, the data format corresponding to the initial object feature in the initial feature set may be the RfcDataObj format, and the data formats supported by the feature database are the JSON format or the binary format, and the RfcDataObj format is not supported. In this case, the feature database cannot store the features in the RfcDataObj format, and it is necessary to convert the initial object features in the RfcDataObj format into candidate object features in the JSON format or the binary format. The specific data format conversion process may include: a format conversion function can be written through a programming language (for example, Python, JavaScript, etc.), and the format conversion function can convert the data in the RfcDataObj format into the data in the JSON format, and then the format conversion function can be called to convert the initial object features in the RfcDataObj format into candidate object features in the JSON format.
[0109] Step S103: Perform feature aggregation processing on the candidate object features in the candidate feature set to obtain object aggregation features, and store the object aggregation features in the feature database.
[0110] The object aggregation feature refers to the feature obtained by aggregating similar candidate object features in the candidate feature set. In the embodiment of the present application, the feature database stores the object aggregation features, and each object aggregation feature requests a storage record from the feature database, without each candidate object feature separately requesting a storage record from the feature database, which can significantly reduce the read and write pressure of the feature database. In addition, when the feature database receives a feature acquisition request corresponding to a service application, it can quickly query the object aggregation feature matching the feature acquisition request from the feature database, which helps to improve the query efficiency of the features.
[0111] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a feature aggregation provided by an embodiment of the present application. As Figure 4 shown, the feature aggregation process may include update (Read&Update), merge (Read&Merge), and output (Serialize&Write). As Figure 4 shown, the data processing node A processes to obtain the initial object feature A, and its corresponding key (KEY) is denoted as KEY1. The data processing node B processes to obtain the initial object feature B, and its corresponding key is denoted as KEY2. As Figure 4As shown, data processing node A reads the data formats supported by the feature database from the service configuration file, converts the data format of the initial object feature A1 to obtain the candidate object feature A1. The data format corresponding to the candidate object feature A1 is the data format supported by the feature database, for example, the JSON format. Similarly, data processing node B reads the data formats supported by the feature database from the service configuration file, converts the data format of the initial object feature B1 to obtain the candidate object feature B1. The data format corresponding to the candidate object feature B1 is the data format supported by the feature database, for example, the JSON format. Optionally, in order to try to stagger the update time intervals of different data processing nodes, a delay function (sleep) can be used to add a delay of 0 - 100 ms to the data processing nodes before the update, so that the generation timestamps (timestamps) of the candidate object features output by different data processing nodes are different, which is convenient for subsequent offline stitching and at the same time alleviates the problem of write coverage.
[0112] When the aggregation period arrives, data processing node A can obtain the candidate object feature with the key KEY1 (for example, candidate object feature A2) from other processing nodes (for example, data node processing node B), merge all the candidate object features with the key KEY1 (candidate object feature A1 and candidate object feature A2) to obtain the object combination feature A, and then perform serialization processing on the object combination feature A to obtain the object aggregation feature A. Similarly, data processing node B can obtain the candidate object feature with the key KEY2 (for example, candidate object feature B2) from other processing nodes (for example, data node processing node B), merge all the candidate object features with the key KEY2 (candidate object feature B1 and candidate object feature B2) to obtain the object combination feature B, and then perform serialization processing on the object combination feature B to obtain the object aggregation feature B. Serialization processing can refer to the process of converting the data format of the feature into a binary format. Serialization tools such as Flat Buffers and ProtocolBuffers (a serialization tool) can be used to serialize the object combination feature into the object aggregation feature.
[0113] In a possible implementation manner, feature aggregation can be performed based on object identification information, which specifically may include: obtaining the object identification information corresponding to the candidate object features in the candidate feature set, combining the candidate object features with the same object identification information to obtain the first combination feature, and generating the object aggregation feature with the object identification information as the key and the first combination feature as the value.
[0114] Please refer to Figure 5 , Figure 5 which is a schematic diagram of a feature aggregation provided by an embodiment of the present application. As Figure 5As shown, the candidate feature set may include candidate object feature 40a, candidate object feature 40b, candidate object feature 40c, and candidate object feature 40d. Among them, the object identification information corresponding to candidate object feature 40a, candidate object feature 40b, candidate object feature 40c, and candidate object feature 40d is all the object identification information (user_A_id) of business object A. The candidate object feature 40a, candidate object feature 40b, candidate object feature 40c, and candidate object feature 40d can be combined to obtain the first combined feature 40e; then, using the object identification information of business object A as the key and the first combined feature 40e as the value, generate as Figure 5 the object aggregation feature 40f shown.
[0115] Optionally, after obtaining the first combined feature, the first combined feature can also be serialized to obtain the first serialized feature corresponding to the first combined feature, and then, using the object identification information as the key and the first serialized feature as the value, generate the object aggregation feature. Serialization can make the online reading and writing efficiency of the object aggregation feature more efficient. The process of serialization can refer to the above description and will not be elaborated here.
[0116] In a possible implementation manner, feature aggregation can be performed based on key fields, specifically including: obtaining the key fields corresponding to the candidate object features in the candidate feature set, combining the candidate object features with the same key fields to obtain a second combined feature, and using the key field as the key and the second combined feature as the value to generate the object aggregation feature.
[0117] Among them, it refers to the field in the field information corresponding to the candidate object feature that represents the key information. The key field can be set according to the actual business situation. For example, if the field information corresponding to the candidate object feature is "skin_buy_seq_1h (indicating the skin identifier for transactions within 1 hour continuously)", then the key field corresponding to this candidate object feature can be "skin_buy". For example, the candidate feature set may include candidate object feature 50a, candidate object feature 50b, candidate object feature 50c, and candidate object feature 50d. The key fields corresponding to candidate object feature 50a and candidate object feature 50b are "skin_buy", and the key fields corresponding to candidate object feature 50c and candidate object feature 50d are "character_buy". Then, combine candidate object feature 50a and candidate object feature 50b to obtain the second combined feature 50e. Using the key field "skin_buy" as the key and the second combined feature 50e as the value, generate the object aggregation feature 50g; similarly, candidate object feature 50c and candidate object feature 50d can be combined to obtain the second combined feature 50f. Using the key field "character_buy" as the key and the second combined feature 50f as the value, generate the object aggregation feature 50h.
[0118] Optionally, after obtaining the second combined feature, the second combined feature can be serialized to obtain the second serialized feature corresponding to the second combined feature. Using the key field as the key and the second serialized feature as the value, generate the object aggregation feature. Serialization processing can make the online reading and writing efficiency of the object aggregation feature more efficient. The process of serialization processing can refer to the above description and will not be elaborated here.
[0119] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a feature storage provided by an embodiment of the present application. Figure 1 As Figure 6 shown, after obtaining the object aggregation feature, the feature production component can store the object aggregation feature in the feature database to provide a data source for the business recommendation service corresponding to the business application. The specific business recommendation process will be described in detail below and will not be elaborated here.
[0120] Optionally, to maintain data consistency, the features before serialization processing (e.g., the first combined feature and the second combined feature mentioned above) can be stored to maintain data consistency and also facilitate business recommendation services that do not require serialized features. To reduce data storage pressure, the features before serialization processing and the features after serialization processing can be stored in different databases respectively. For example, the features before serialization processing can be stored in the cache database mentioned below, and the features after serialization processing can be stored in the feature database.
[0121] Optionally, after obtaining the object aggregation feature, the object aggregation feature can also be subjected to feature compression processing to obtain the compressed object aggregation feature, and then the compressed object aggregation feature is stored in the feature database to improve the read and write performance. The compression algorithms used in the embodiments of the present application can include, but are not limited to, one or more of the compression algorithms such as Huffman compression algorithm, LZW compression algorithm, ZSTD compression algorithm, etc. The string length after the object aggregation feature is compressed can be shortened to 1 / 3 after compression, which helps to improve the feature read and write performance of the online service of the business recommendation system.
[0122] In the embodiments of the present application, the feature database supported by the business application can be used to store the object aggregation features produced by the feature production component for the business application, so as to realize the transplantation and access of the feature production component into the business recommendation system in the business application. Since the feature database is the database supported by the business application, the feature database and the business recommendation system in the business application have good environmental adaptability, which can simplify the component dependency between the feature production component and the business recommendation system, and there is no need to perform secondary development on the feature generation component, reducing the access cost of the feature production component. In addition, the feature production component can be applied to multiple business applications that support the feature database without independently developing the feature generation component according to the architecture of the business recommendation system, improving the versatility of the feature production component.
[0123] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of a data processing method provided by the embodiments of the present application Figure 2 and this data processing method can be executed by a computer device, which can be Figure 1 the terminal device 10a shown in Figure 1 or it can also be Figure 7 the server 10d shown in
[0124] Step S201: Obtain the initial feature set corresponding to the business application through the feature production component, and determine the feature database corresponding to the business application according to the database parameters included in the business configuration file in the feature production component.
[0125] Step S202: Perform data format conversion on the initial object features in the initial feature set to obtain the candidate feature set.
[0126] Among them, the specific implementation processes of Step S201 and Step S202 can refer to Figure 3 the descriptions of Step S101 and Step S102 shown, and will not be elaborated here.
[0127] Step S203: If the business configuration file contains the database parameters for caching the candidate feature set, determine the cache database corresponding to the feature production component.
[0128] It can be understood that the feature production component will frequently read and write to the feature database. Each time a feature is stored in the feature database, the number of operations required is: read times (read_num) * (1 + merge times (merge_num)) + write times (write_num) * 2. It can be seen that the feature database has a large read and write pressure, and thus may easily cause problems such as too high read and write latency and write overwrite when providing data sources to business applications.
[0129] The embodiment of the present application can adopt the method of storage separation, that is, the candidate object features and the object aggregation features are stored in different databases respectively. Specifically, please refer to Figure 8 , Figure 8 which is a schematic diagram of a feature storage provided by the embodiment of the present application Figure 2 . As Figure 8 shown, the candidate object features can be stored in the cache database waiting for subsequent aggregation operations, and the aggregated object aggregation features are stored in the feature database for business applications to use.
[0130] It can be understood that the expiration time of the candidate object features generated by the feature production component is usually very short (<1d), and the capacity requirement for the cache database is not high. Therefore, a database with high read and write performance and small capacity (<30G) can be selected as the cache database, such as the Cache Redis database (a lightweight key-value pair database), the LevelDB database (a lightweight key-value pair database), etc. The feature database directly interacts with the business application, so a database with high read and write performance and high capacity can be selected as the cache database, such as the SSD Redis database (a high-capacity key-value pair database), the NoSQL database, etc.
[0131] The implementation process of the cache database access feature production component can refer to Figure 3 the relevant description of step S101 shown above, which will not be elaborated here. When connecting the cache database to the feature production component, the database parameters corresponding to the cache database can be written into the service configuration file. Therefore, it is possible to determine whether the cache database is connected to the feature production component by checking the service configuration file.
[0132] If the service configuration file in the feature production component contains not only the database parameters for caching object aggregation features but also the database parameters for caching the candidate feature set, then according to these database parameters, the cache database corresponding to the feature production component can be determined, and then the candidate feature set can be stored in this cache database.
[0133] Step S204: In the cache database, determine the candidate feature set whose generation timestamp belongs to the aggregation period as the to-be-processed feature set.
[0134] In the embodiment of the present application, the number of candidate feature sets cached in the cache database is N, and a candidate feature set carries a generation timestamp. N is a positive integer, and the specific value of N can be 1, 2, 5, 10, etc. The generation timestamp can refer to the timestamp when the initial feature set corresponding to the candidate feature set is output from the data processing node. The aggregation period refers to the time interval composed of the start time and the end time of feature aggregation. It is possible to search in the cache database for the candidate feature sets whose generation timestamps are within the aggregation period, and determine the candidate feature sets whose generation timestamps are within the aggregation period as the to-be-processed feature sets.
[0135] For example, the start time corresponding to the aggregation period is "20xx - 10:00", and the end time is "20xx - 11:00". The cache database stores candidate feature set a, candidate feature set b, and candidate feature set c. The generation timestamp corresponding to candidate feature set a is "20xx - 10:05", the generation timestamp corresponding to candidate feature set b is "20xx - 10:15", and the generation timestamp corresponding to candidate feature set c is "20xx - 09:57". At this time, it can be determined that the generation timestamps corresponding to candidate feature set a and candidate feature set b belong to the aggregation period, and candidate feature set a and candidate feature set b can be determined as the to-be-processed feature sets.
[0136] Step S205: Perform feature aggregation processing on the candidate object features in the to-be-processed feature set to obtain object aggregation features, store the object aggregation features in the feature database, and clear the to-be-processed feature set in the cache database.
[0137] After obtaining the set of features to be processed, the candidate object features in the set of features to be processed can be subjected to feature aggregation processing to obtain object aggregation features; after the aggregation is completed, the set of features to be processed can be cleared in the cache database to save the data storage space of the cache database. The specific feature aggregation method can refer to Figure 3 the relevant description of step S103 shown, which will not be elaborated here. It can be seen that the embodiments of the present application can aggregate the candidate object features in the candidate feature sets with different generation timestamps, which can reduce the time for feature aggregation and thus improve the efficiency of feature aggregation.
[0138] Optionally, after obtaining the object aggregation features, the object aggregation features can be subjected to feature compression processing to obtain object encoding features, and the object encoding features are stored in the feature database, thereby further improving the read and write performance. The compression process of the object aggregation features can include: obtaining D unit characters included in the object aggregation features, counting the occurrence frequencies of each unit character in the object aggregation features, and determining the occurrence frequencies corresponding to each unit character as the weight values corresponding to each unit character. Where D represents the number of unit characters included in the object aggregation features, D is an integer greater than 1, and the specific value of D can be 2, 5, 10, etc.; the unit character refers to the characters without repetition in the object aggregation features. For example, if the string included in the object aggregation features is "skin_id: 005", the unit characters can include: "s", "k", "i", "n", "_", "d", ":", "0", "5"; the occurrence frequencies of the unit characters "s", "k", "n", ":", "_", "d", and "5" are all 1, and the occurrence frequencies of the unit characters "i" and "0" are 2. The weight values corresponding to the unit characters "s", "k", "n", ":", "_", "d", and "5" can be set to 1, and the weight values corresponding to the unit characters "i" and "0" can be set to 2; optionally, the weight values corresponding to the unit characters "s", "k", "n", ":", "_", "d", and "5" can also be set to 0.1, and the weight values corresponding to the unit characters "i" and "0" can be set to 0.2. The specific values of the weight values corresponding to each unit character can be set according to the actual situation.
[0139] Further, a feature encoding tree corresponding to the object aggregation feature can be constructed according to each unit character and the weight value corresponding to each unit character. The embodiments of the present application do not limit the construction method of the feature encoding tree, nor do they limit the shape and depth of the feature encoding tree. The feature encoding tree has the following characteristics: ① The leaf nodes in the feature encoding tree are the nodes corresponding to D unit characters, that is, the leaf nodes in the feature encoding tree can be used to represent D unit characters; ② The weight value corresponding to the parent node i in the feature encoding tree is the sum of the weight values corresponding to the child nodes of the parent node i. The parent node i is any parent node in the feature encoding tree. For example, the child nodes of the parent node i are the leaf nodes corresponding to the unit character "s" and the leaf nodes corresponding to the unit character "k", and the weight values corresponding to the unit character "s" and the unit character "k" are both 1, then the weight value corresponding to the parent node i is 2.
[0140] After the feature encoding tree is constructed, encoding information can be set for the nodes in the feature encoding tree. For example, the path to the left in the feature encoding tree is marked as "0", and the path to the right is marked as "1", or other symbols can also be used to replace "0" and "1"; then the shortest paths between each leaf node and the root node in the feature encoding tree are obtained, and according to the position information of each node in the shortest path in the feature encoding tree, the encoding information of the unit character corresponding to each leaf node is obtained. That is to say, in the feature encoding tree, the unit characters with low occurrence frequency (small weight value) are given longer encodings, while the unit characters with high occurrence frequency (large weight value) are given shorter encodings. Assume that the shortest path between the leaf node c and the root node a is {root node a → node b → leaf node c}, and node b is the left node of the feature encoding tree, and leaf node c is the right node of the feature encoding tree, then the encoding information of the unit character corresponding to leaf node c can be expressed as "01". Similarly, the encoding information corresponding to the D unit characters in the object aggregation feature can be obtained in the same way.
[0141] Further, the object encoding feature corresponding to the object aggregation feature can be generated according to the encoding information corresponding to each unit character and the text position of each unit character in the object aggregation feature; for example, the string included in the object aggregation feature is "aabcba", the encoding information corresponding to the unit character "a" is "0", the encoding information corresponding to the unit character "b" is "10", and the encoding information corresponding to the unit character "c" is "11", then the object encoding feature corresponding to the object aggregation feature can be expressed as "001011100". After obtaining the object encoding feature corresponding to the object aggregation feature, the object encoding feature can be stored in the feature database to provide a data source for the service recommendation service corresponding to the business application.
[0142] In the embodiments of the present application, a feature coding tree with each unit character in the object aggregation feature as leaf nodes is constructed according to the occurrence frequency of each unit character in the object aggregation feature; then, according to the position information of each node in the shortest path in the feature coding tree, the coding information of the unit character corresponding to each leaf node is determined, and the object coding feature is generated according to the coding information of each unit character. Among them, the weight value corresponding to the parent node in the feature coding tree is the sum of the weight values corresponding to the child nodes of the parent node; this means that during the construction of the feature coding tree, the unit character with a small weight value (low occurrence frequency) is assigned a longer code, while the unit character with a large weight value (high occurrence frequency) is assigned a shorter code. This coding method can intuitively and quickly determine the coding information corresponding to each unit character through the feature coding tree, and can significantly reduce the average length of the strings included in the object aggregation feature after coding while ensuring data integrity, thereby reducing the data storage space and transmission time.
[0143] Please refer to Figure 9 , Figure 9 which is a schematic diagram of a feature processing provided by the embodiments of the present application. As Figure 9 shown, the feature production component can obtain an initial feature set through a data processing node, perform data format conversion on the initial object features in the initial feature set to obtain a candidate feature set; then perform feature aggregation processing on the candidate object features in the candidate feature set to obtain an object aggregation feature, and perform feature compression processing on the object aggregation feature to obtain an object coding feature. Among them, the specific implementation processes of format conversion, feature aggregation, and feature compression can refer to the above description and will not be elaborated here.
[0144] It can be understood that the business recommendation service in the business application can include online services and can also include offline services. The object coding feature can be a real-time feature, and the object coding feature can be published to the feature database to provide a data source for online services. For example, real-time business recommendations can be made according to the object coding feature. Optionally, in order to enable the offline samples in the business recommendation system to include real-time features, the real-time features also need to be stored to provide a data source for offline services. For example, offline model training can be performed according to the object coding feature, etc. If the business application enables the feature snapshot function, the object coding feature can be stored in the feature snapshot library in the business recommendation system in the form of a feature snapshot (such as Figure 2The feature snapshot library 207) shown. If the feature snapshot function is not enabled in the business application, the object-encoded features can be stored in the data collection system (e.g., TDBank) and then stored in the distributed data warehouse (e.g., TDW) in the business application to form sample features, which are offline real-time features; then, the sample features and sample labels are spliced and sampled to generate training samples, which is convenient for providing data sources for offline services. It can be seen that the feature publishing method in the embodiments of the present application can well guarantee the service quality of both the online service and the offline service of the business recommendation system.
[0145] Step S206: If a business recommendation request is received, obtain the object identification information corresponding to the business recommendation request.
[0146] Step S207: Determine the object recommendation features corresponding to the business recommendation request as the object aggregation features in the feature database that match the object identification information.
[0147] For ease of understanding, in the embodiments of the present application, a virtual skin recommendation scenario in a game application (e.g., Game A) is taken as an example to describe the specific process of business recommendation. The game application can be any one of third-person shooting (TPS) games, first-person shooting (FPS) games, multiplayer online battle arena games (MOBA), massive multiplayer online role-playing games (MMORPG), strategy games, etc.
[0148] Specifically, please refer to Figure 10 , Figure 10 which is a schematic diagram of an interface for recommending virtual skins provided by the embodiments of the present application. As Figure 10 shown, the business object can perform a trigger operation on the login control of Game A, and Game A responds to the trigger operation on the login control to display the login page 30a of Game A. The business object can enter the account information and password information in the login information input area 30b on the login page 30a and perform a trigger operation on the login game control 30c; Game A responds to the trigger operation on the game control 30c, verifies the input information in the login information input area 30b, and after the verification passes, displays the home page 31a of Game A.
[0149] Generally speaking, the residence time of a business object on the game progress page in Game A is much longer than that on the game mall page 33a. That is to say, before entering the game mall page 33a, the business object has often performed interactive operations such as battles, matching rankings, or team operations, and viewing the items held in the backpack (such as virtual characters, virtual skins, and virtual props) in Game A. At this time, the interaction data generated by the above interactive operations can be used as object metadata and passed into the data processing node in the feature production component to obtain real-time object aggregation features. Furthermore, item recommendations can be made based on the above real-time object aggregation features, which can match more items that conform to the interest tendency of the business object and help improve the accuracy of item recommendations.
[0150] As Figure 10 shown, the business object can perform a trigger operation on the game mall entrance 31b on the home page 31a. Game A responds to the trigger operation for the game mall entrance 31b, generates a business recommendation request corresponding to the business object, and sends the business recommendation request to the background server 34a corresponding to Game A. The background server 34a receives the business recommendation request and sends a feature acquisition request to the feature database 34b. The feature acquisition request can carry the object identification information corresponding to the business recommendation request, and requests to obtain the object aggregation feature matching the object identification information from the feature database 34b through the feature acquisition request, and determines the object aggregation feature as the object recommendation feature corresponding to the business recommendation request. Optionally, the business recommendation request can also carry a feature acquisition time range, and the object aggregation feature in the feature database 34b can be further filtered according to the feature acquisition time range to improve the accuracy of item recommendations. For example, the feature acquisition time range can be 1 hour. The object aggregation features with a generation timestamp not exceeding 1 hour can be filtered out from the object aggregation features matching the object identification information and determined as the object recommendation features corresponding to the business recommendation request, thereby improving the real-time nature of the object recommendation features.
[0151] Step S208: Obtain the set of to-be-recommended services corresponding to the business recommendation request, and obtain the business recommendation features corresponding to each to-be-recommended service in the set of to-be-recommended services.
[0152] Step S209: Obtain the similarity evaluation values between the business recommendation features corresponding to each to-be-recommended service and the object recommendation feature, and determine the recommendation status corresponding to each to-be-recommended service according to the similarity evaluation values.
[0153] Step S210: Display the to-be-recommended services whose recommendation status belongs to the recommended start status in the business application.
[0154] It can be understood that some to-be-recommended services (such as virtual characters, virtual skins, virtual props, etc.) can be screened out from the item recommendation library and added to the set of to-be-recommended services. For example, to-be-recommended services with higher exposure or transaction volume can be screened out from the item recommendation library to form the set of to-be-recommended services corresponding to the service recommendation request, so as to improve the accuracy of item recommendation. Furthermore, business recommendation features corresponding to each to-be-recommended service in the set of to-be-recommended services can be obtained, and the business recommendation features can be used to characterize the attribute information of the to-be-recommended services.
[0155] Furthermore, a similarity evaluation value between the business recommendation features corresponding to each to-be-recommended service and the object recommendation features can be obtained. The similarity evaluation value can be understood as the similarity between the business recommendation features and the object recommendation features. The similarity between the business recommendation features and the object recommendation features can be characterized by calculating the distance between the business recommendation features and the object recommendation features. The shorter the distance between the business recommendation features and the object recommendation features, the greater the similarity between the two; the greater the distance between the business recommendation features and the object recommendation features, the smaller the similarity between the two. Among them, the methods for calculating the distance between the business recommendation features and the object recommendation features can include but are not limited to: Euclidean distance, Manhattan distance, Minkowski distance, Cosine Similarity, etc. Optionally, the similarity evaluation value can be numerically displayed in the form of a numerical value. For example, the similarity evaluation value can be a numerical value between 0 and 1, or it can also be a numerical value between 0 and 100. The larger the numerical value, the greater the similarity evaluation value, and the greater the similarity evaluation value between the business recommendation features and the object recommendation features.
[0156] Optionally, the business recommendation features and object recommendation features corresponding to each business to be recommended can also be input into a business recommendation model. Through the feature extraction layer in the business recommendation model, feature extraction is performed on the business recommendation features corresponding to each business to be recommended to obtain business embedding features corresponding to each business recommendation feature; through the feature extraction layer in the business recommendation model, feature extraction is performed on the object recommendation feature to obtain object embedding features corresponding to the object recommendation feature; and then through the matching layer in the business recommendation model, a similarity evaluation value between each business embedding feature and the object embedding feature is output. The business recommendation model can be a two-tower model (Deep Structured Semantic Models, DSSM), a deep neural network model (Deep Neural Networks, DNN), a convolutional neural network model (Convolutional Neural Networks, CNN), a graph neural network model, etc., or can be a variant of any of the above models; the network structure of the feature extraction layer in the business recommendation model can include but is not limited to a DNN network, a CNN network, etc., or can be a variant of any of the above networks; the matching layer of the business recommendation model can include one or more of the above similarity algorithms.
[0157] Furthermore, the recommendation status corresponding to each business to be recommended can be determined according to the similarity evaluation value. Generally speaking, the larger the similarity evaluation value, the greater the recommendation probability of the corresponding business to be recommended. For ease of description, taking the number of businesses to be recommended in the set of businesses to be recommended as L as an example, the determination process of the recommendation status corresponding to each business to be recommended is described; where L is a positive integer, and the specific value of L can be 5, 10, 15, etc.
[0158] Specifically, the similarity evaluation values corresponding to the L businesses to be recommended can be sorted in descending order to obtain the sorted L businesses to be recommended; for example, the L businesses to be recommended are: Business to be Recommended 1, Business to be Recommended 2, Business to be Recommended 3, Business to be Recommended 4, and Business to be Recommended 5. The similarity evaluation value corresponding to Business to be Recommended 1 is 0.4, the similarity evaluation value corresponding to Business to be Recommended 2 is 0.2, the similarity evaluation value corresponding to Business to be Recommended 3 is 0.8, the similarity evaluation value corresponding to Business to be Recommended 4 is 0.9, and the similarity evaluation value corresponding to Business to be Recommended 5 is 0.7. The sorted L businesses to be recommended are {Business to be Recommended 4, Business to be Recommended 3, Business to be Recommended 5, Business to be Recommended 1, Business to be Recommended 2}.
[0159] Furthermore, obtain the number y (y is a positive integer less than or equal to L) corresponding to the recommendation display position in the business application; the recommendation display position refers to the area in the business application for displaying the business to be recommended. For example, in Figure 10Among them, the recommended display positions included in the skin recommendation area 33g on the game mall page 33a include recommended display position 33c, recommended display position 33d, recommended display position 33e, and recommended display position 33f, and the number y of the recommended display positions included in the skin recommendation area 33g is 4. Among the sorted L to-be-recommended services, the recommended status corresponding to the first y to-be-recommended services is determined as the recommended enabled status, and the recommended status corresponding to the to-be-recommended services other than the first y to-be-recommended services is determined as the recommended disabled status. For example, the recommended status corresponding to to-be-recommended service 4, to-be-recommended service 3, to-be-recommended service 5, and to-be-recommended service 1 is determined as the recommended enabled status, and the recommended status corresponding to to-be-recommended service 2 is determined as the recommended disabled status; and then the to-be-recommended services with the recommended status belonging to the recommended enabled status are displayed in the service application.
[0160] As Figure 10 shown, after determining the recommended status corresponding to each to-be-recommended information, the background server 34a may send the to-be-recommended services with the recommended status belonging to the recommended enabled status to the terminal device and display the to-be-recommended service on the game mall page 33a. As Figure 10 shown, the service object may perform a trigger operation on the skin tab 33b, and game A responds to the skin tab 33b and displays the virtual skins with the recommended status belonging to the recommended enabled status in the recommended display positions in the skin recommendation area 33g. For example, the recommended status of virtual skin one, virtual skin two, virtual skin three, and virtual skin four is the recommended enabled status, virtual skin one can be displayed in the recommended display position 33c, virtual skin two can be displayed in the recommended display position 33d, virtual skin three can be displayed in the recommended display position 33e, and virtual skin four can be displayed in the recommended display position 33f. Alternatively, the service object may also perform a trigger operation on the prop tab in the game mall page 33a, and game A responds to the prop tab and displays the virtual props with the recommended status belonging to the recommended enabled status in the recommended display positions in the prop recommendation area, etc.
[0161] To verify the effect of the transplantation and access scheme of the feature production component provided in the embodiment of the present application, a feature production component may be accessed in the service recommendation system of game A. Among them, Figure 11a to Figure 11c it feedbacks the overall delay situation during the feature production process. Specifically, please refer to Figure 11a , Figure 11a which is the schematic Figure 1 of the feature read / write delay provided in the embodiment of the present application. From Figure 11a it can be seen that when the feature production component produces real-time features for game A, the delay generated by the real-time features themselves is approximately 6 seconds - 10 seconds. Please refer to Figure 11b , Figure 11b which is the schematic Figure 2 of the feature read / write delay provided in the embodiment of the present application. From Figure 11bIt can be seen that the feature production component can process 100,000 object metadata of Game A per minute on average. Please refer to Figure 11c , Figure 11c which is a schematic diagram of the feature read-write delay provided by an embodiment of the present application. Figure 3 From Figure 11c it can be seen that the process of the feature production component producing real-time features for Game A takes approximately 200 ms - 400 ms. When the feature production component is transplanted and connected to the business recommendation system of Game A in the embodiment of the present application, it can fully meet the business requirements of Game A.
[0162] Please refer to Figure 12 , Figure 12 which is a schematic diagram of the feature database read delay provided by an embodiment of the present application. The feature storage scheme before optimization is the feature storage scheme as shown in Figure 6 , and the optimized feature storage scheme is the feature storage scheme as shown in Figure 8 . Using P999 as the evaluation index for both, where P999 means that 99.9% of the data is less than or equal to the threshold. As shown in Figure 12 , when the optimized feature storage scheme is adopted, when accessing the feature database online, the access delay P999 drops from 500 ms during peak hours to 10 ms, greatly improving the service quality.
[0163] Since the object aggregation feature is the feature after the candidate object features are feature-aggregated, and the feature aggregation operation is prone to bring about the write-overwrite problem, therefore, in terms of feature quality, the main concern is the integrity of the features. Among them, the write-overwrite problem refers to the situation where, among the features to be aggregated, when a certain feature is written into the feature database, it overwrites the feature values of other features, resulting in other features not being updated. The write-overwrite situation is shown in Table 1 below:
[0164] Table 1
[0165]
[0166] As shown in Table 1, the features in the second row of Table 1 overwrite the features corresponding to Seq_1h and Seq_6h produced by the first row, and finally the feature only records the features corresponding to Cnt_1h and Cnt_6h. The write-overwrite problem is usually because the features to be aggregated read various old data, and when updating their respective data and writing, the later-generated features will overwrite the updates of other features. A too high write-overwrite rate will seriously affect the accuracy of the features, resulting in the situation where the features are unusable. After adopting the access and transplantation scheme of the embodiment of the present application, the write coverage rate drops from 98.2% to 0.6%, basically solving the write-overwrite problem, so that the quality of the features produced by the feature production component can be effectively guaranteed.
[0167] Further, after the transplantation and access are completed, the business recommendation effect in Game A is evaluated. The AUC is used as the evaluation index for the business recommendation model. The higher the AUC, the better the prediction effect of the business recommendation model. The data of the current day can be used as the validation set, and the data before the previous day can be used as the training set. Specifically, please refer to Figure 13 , Figure 13 which is a schematic diagram of the business recommendation effect provided by an embodiment of the present application. As Figure 13 shown, the prediction effect of the business recommendation model trained offline using real-time features is significantly better than that of the business recommendation model trained offline without using real-time features. The specific prediction effect is shown in Table 2 below:
[0168] Table 2
[0169]
[0170] As shown in Table 2, for the business recommendation model trained using real-time features, compared with the business recommendation model trained without using real-time features, the overall offline AUC of the test set has increased by 1.03%, the online AUC has increased by 1.49%, and the online GAUC has increased by 0.91%.
[0171] In the embodiment of the present application, the feature database supported by the business application can be used to store the object aggregation features generated by the feature production component for the business application, so as to realize the transplantation and access of the feature production component into the business recommendation system in the business application. Since the feature database is the database supported by the business application, the feature database and the business recommendation system in the business application have good environmental adaptability, which can simplify the component dependency between the feature production component and the business recommendation system, without the need for secondary development of the feature generation component, and reduce the access cost of the feature production component. In addition, the feature production component can be applied to multiple business applications that support the feature database, without the need for independent development of the feature generation component according to the architecture of the business recommendation system, improving the versatility of the feature production component.
[0172] It can be understood that in the specific implementation manner of the present application, it may involve relevant information of users (for example, user account information, user password information, etc.). When the above embodiments of the present application are applied to specific products or technologies, the permission or consent of the users needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards in the relevant regions.
[0173] Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. It can be understood that the data processing device 1 can be applied to a computer device (such as Figure 1in the terminal device or server shown. As Figure 14 As shown, the data processing device 1 may include: a database determination module 11, a format conversion module 12, and a feature aggregation module 13:
[0174] The database determination module 11 is configured to obtain an initial feature set corresponding to a service application through a feature production component, and determine a feature database corresponding to the service application according to database parameters included in a service configuration file in the feature production component;
[0175] The format conversion module 12 is configured to perform data format conversion on initial object features in the initial feature set to obtain a candidate feature set; the data format corresponding to the candidate object features in the candidate feature set is the data format supported by the feature database;
[0176] The feature aggregation module 13 is configured to perform feature aggregation processing on the candidate object features in the candidate feature set to obtain an object aggregation feature, and store the object aggregation feature in the feature database; the feature database is used to provide a data source for a service recommendation service corresponding to the service application.
[0177] In a possible implementation manner, the database determination module 11 is specifically configured to:
[0178] Call a service data query interface to obtain an object metadata set from a service database corresponding to the service application;
[0179] Batch-process the object metadata in the object metadata set to obtain M data subsets; M is an integer greater than 1;
[0180] Distribute the M data subsets to M data processing nodes in the feature production component, and perform parallel processing on the object metadata in the M data subsets through the M data processing nodes to obtain object output features corresponding to the M data processing nodes; the object metadata in one data subset is used to be distributed to one data processing node;
[0181] Merge the object output features corresponding to the M data processing nodes to obtain an initial feature set corresponding to the service application.
[0182] In a possible implementation manner, the feature aggregation module 13 is further configured to:
[0183] If the service configuration file includes database parameters for caching the candidate feature set, determine a cache database corresponding to the feature production component; the number of candidate feature sets cached in the cache database is N, and one candidate feature set carries a generation timestamp, and N is a positive integer;
[0184] In the cache database, determine the candidate feature set whose generation timestamp belongs to the aggregation period as the to-be-processed feature set;
[0185] Perform feature aggregation processing on the candidate object features in the to-be-processed feature set to obtain object aggregation features, and clear the to-be-processed feature set in the cache database.
[0186] In a possible implementation manner, the feature aggregation module 13 is specifically configured to:
[0187] Obtain the object identification information corresponding to the candidate object features in the candidate feature set, combine the candidate object features with the same object identification information to obtain a first combined feature;
[0188] Perform serialization processing on the first combined feature to obtain a first serialized feature corresponding to the first combined feature, and generate object aggregation features with the object identification information as the key and the first serialized feature as the value.
[0189] In a possible implementation manner, the feature aggregation module 13 is specifically configured to:
[0190] Obtain the key fields corresponding to the candidate object features in the candidate feature set, combine the candidate object features with the same key fields to obtain a second combined feature;
[0191] Perform serialization processing on the second combined feature to obtain a second serialized feature corresponding to the second combined feature, and generate object aggregation features with the key fields as the key and the second serialized feature as the value.
[0192] In a possible implementation manner, the feature aggregation module 13 is specifically configured to:
[0193] Obtain D unit characters included in the object aggregation features, count the occurrence frequencies of each unit character in the object aggregation features, and determine the occurrence frequencies corresponding to each unit character as the weight values corresponding to each unit character; D is an integer greater than 1;
[0194] Construct a feature coding tree corresponding to the object aggregation features according to each unit character and the weight value corresponding to each unit character; the leaf nodes in the feature coding tree are used to represent D unit characters; the weight value corresponding to the parent node i in the feature coding tree is the sum of the weight values corresponding to the child nodes of the parent node i;
[0195] Obtain the shortest paths between each leaf node and the root node in the feature coding tree, and obtain the coding information of the unit characters corresponding to each leaf node according to the position information of each node in the shortest path in the feature coding tree;
[0196] Generate an object coding feature corresponding to the object aggregation feature according to the coding information corresponding to each unit character and the text position of each unit character in the object aggregation feature, and store the object coding feature in the feature database.
[0197] In a possible implementation, the data processing device further includes a database access module 14, and the database access module 14 is configured to:
[0198] Create a data storage class for inheriting the database access interface in the feature generation component, and create a data storage function in the data storage class;
[0199] Determine the database identification information and data storage structure corresponding to the feature database as the database parameters corresponding to the feature database;
[0200] Generate the implementation logic of the data storage function according to the database access interface and the database parameters corresponding to the feature database;
[0201] Execute the implementation logic of the data storage function, and write the database parameters corresponding to the feature database into the service configuration file.
[0202] In a possible implementation, the data processing device further includes a service recommendation module 15, and the service recommendation module 15 is configured to:
[0203] If a service recommendation request is received, obtain the object identification information corresponding to the service recommendation request;
[0204] Determine the object aggregation feature in the feature database that matches the object identification information as the object recommendation feature corresponding to the service recommendation request;
[0205] Obtain the set of services to be recommended corresponding to the service recommendation request, and obtain the service recommendation features corresponding to each service to be recommended in the set of services to be recommended;
[0206] Obtain the similarity evaluation value between the service recommendation feature corresponding to each service to be recommended and the object recommendation feature, and determine the recommendation status corresponding to each service to be recommended according to the similarity evaluation value;
[0207] Display the services to be recommended whose recommendation status belongs to the recommended start status in the service application.
[0208] In a possible implementation, the number of services to be recommended in the set of services to be recommended is L, and L is a positive integer; the service recommendation module 15 is specifically configured to:
[0209] Sort the similarity evaluation values corresponding to the L services to be recommended in descending order to obtain the sorted L services to be recommended;
[0210] Obtain the quantity y corresponding to the recommended display position in the business application. Among the L to-be-recommended businesses after sorting, determine the recommended status of the first y to-be-recommended businesses as the recommended enabled status, and determine the recommended status of the to-be-recommended businesses other than the first y to-be-recommended businesses as the recommended disabled status; y is a positive integer less than or equal to L.
[0211] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0212] According to an embodiment of the present application, the Figure 3 and Figure 7 steps involved in the data processing method shown can be Figure 14 executed by each module in the data processing device 1 shown. For example, Figure 3 the step S101 shown can be executed by Figure 14 the database determination module 11 shown, Figure 3 the step S102 shown can be executed by Figure 14 the format conversion module 12 shown, Figure 3 the step S103 shown can be executed by Figure 14 the feature aggregation module 13 shown, and so on.
[0213] According to an embodiment of the present application, Figure 14 each module in the data processing device 1 shown can be respectively or entirely combined into one or several units to form, or a certain one (or some) of the units can be further split into at least two smaller sub-units in terms of function, and the same operations can be achieved without affecting the realization of the technical effects of the embodiments of the present application. The above modules are divided based on logical functions. In actual applications, the function of one module can also be implemented by at least two units, or the functions of at least two modules can be implemented by one unit. In other embodiments of the present application, the data processing device 1 can also include other units. In actual applications, these functions can also be assisted by other units and can be implemented by the cooperation of at least two units.
[0214] In the embodiments of the present application, the feature database supported by the business application can be used to store the object aggregation features generated by the feature production component for the business application, so as to implement the transplantation and access of the feature production component into the business recommendation system in the business application. Since the feature database is the database supported by the business application, the feature database and the business recommendation system in the business application have good environmental adaptability, which can simplify the component dependencies between the feature production component and the business recommendation system, without the need for secondary development of the feature generation component, and reduce the access cost of the feature production component. In addition, the feature production component can be applied to multiple business applications that support the feature database, without the need for independent development of the feature generation component according to the architecture of the business recommendation system, improving the versatility of the feature production component.
[0215] Please refer to Figure 15 , Figure 15 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device can be a terminal device or a server as shown in Figure 1 . The computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may further include a user interface 1003 and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components; the user interface 1003 may include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally be at least one storage device located far from the aforementioned processor 1001. As shown in Figure 15 , the memory 1005, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.
[0216] Among them, in the Figure 15 shown computer device 1000, the network interface 1004 can provide network communication functions, while the user interface 1003 is mainly used to provide an input interface for users; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to implement:
[0217] Obtaining an initial feature set corresponding to the business application through the feature production component, and determining the feature database corresponding to the business application according to the database parameters included in the business configuration file in the feature production component;
[0218] Convert the data format of the initial object features in the initial feature set to obtain a candidate feature set; the data format corresponding to the candidate object features in the candidate feature set is the data format supported by the feature database.
[0219] Perform feature aggregation processing on the candidate object features in the candidate feature set to obtain object aggregation features, and store the object aggregation features in the feature database; the feature database is used to provide a data source for the business recommendation service corresponding to the business application.
[0220] It should be understood that the computer device 1000 described in the embodiments of the present application can execute the descriptions of the data processing methods in any of the foregoing Figure 3 and Figure 7 corresponding embodiments, and can also execute the descriptions of the data processing device 1 in the foregoing Figure 8 corresponding embodiments, which will not be elaborated here. In addition, the descriptions of the beneficial effects of using the same method will not be elaborated either.
[0221] In addition, it should be pointed out here that: the embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the foregoing data processing device 1, and the computer program includes computer instructions. When the processor executes the computer instructions, it can execute the descriptions of the data processing methods in any of the foregoing Figure 3 and Figure 7 corresponding embodiments. Therefore, it will not be elaborated here. In addition, the descriptions of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the descriptions of the method embodiments of the present application. As an example, the computer instructions can be deployed to be executed on one computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.
[0222] In addition, it should be noted that: the embodiments of the present application also provide a computer program product or a computer program. The computer program product or the computer program may include computer instructions, and the computer instructions may be stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor can execute the computer instructions, so that the computer device executes the foregoing Figure 3 and Figure 7The description of the data processing method in any corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For the technical details not disclosed in the computer program product or computer program embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0223] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0224] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs.
[0225] The modules in the device embodiments of this application can be combined, divided, and deleted according to actual needs.
[0226] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0227] The above-disclosed are only the preferred embodiments of this application. Of course, the scope of rights of this application cannot be limited thereby. Therefore, equivalent changes made according to the claims of this application still fall within the scope covered by this application.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining an initial feature set corresponding to a business application through a feature production component, and determining a feature database corresponding to the business application according to database parameters included in a business configuration file in the feature production component; Performing data format conversion on initial object features in the initial feature set to obtain a candidate feature set; the data format corresponding to candidate object features in the candidate feature set is a data format supported by the feature database; Performing feature aggregation processing on candidate object features in the candidate feature set to obtain object aggregation features, and storing the object aggregation features in the feature database; the feature database is used to provide a data source for a business recommendation service corresponding to the business application.
2. The method according to claim 1, wherein The obtaining an initial feature set corresponding to a business application through a feature production component includes: Invoking a business data query interface to obtain an object metadata set from a business database corresponding to the business application; Batch-processing object metadata in the object metadata set to obtain M data subsets; M is an integer greater than 1; Distributing the M data subsets to M data processing nodes in the feature production component, and performing parallel processing on object metadata in the M data subsets through the M data processing nodes to obtain object output features corresponding to the M data processing nodes; object metadata in one data subset is used to be distributed to one data processing node; Performing merging processing on object output features corresponding to the M data processing nodes to obtain an initial feature set corresponding to the business application.
3. The method according to claim 1, wherein The method further includes: If database parameters for caching a candidate feature set are included in the business configuration file, determining a cache database corresponding to the feature production component; the number of candidate feature sets cached in the cache database is N, one candidate feature set carries a generation timestamp, and N is a positive integer; Then the performing feature aggregation processing on candidate object features in the candidate feature set to obtain object aggregation features includes: In the cache database, determining candidate feature sets whose generation timestamps belong to an aggregation period as a to-be-processed feature set; Performing feature aggregation processing on candidate object features in the to-be-processed feature set to obtain object aggregation features, and clearing the to-be-processed feature set in the cache database.
4. The method according to claim 1, wherein The performing feature aggregation processing on candidate object features in the candidate feature set to obtain object aggregation features includes: Obtaining object identification information corresponding to candidate object features in the candidate feature set, combining candidate object features with the same object identification information to obtain a first combined feature; Performing serialization processing on the first combined feature to obtain a first serialized feature corresponding to the first combined feature, and generating the object aggregation feature with the object identification information as the key and the first serialized feature as the value.
5. The method according to claim 1, wherein The performing feature aggregation processing on candidate object features in the candidate feature set to obtain object aggregation features includes: Obtain the keyword fields corresponding to the candidate object features in the candidate feature set, and combine the candidate object features with the same keyword fields to obtain the second combined feature; Perform serialization processing on the second combined feature to obtain the second serialized feature corresponding to the second combined feature, and generate the object aggregation feature with the keyword field as the key and the second serialized feature as the value.
6. The method according to claim 1, wherein The storing the object aggregation feature in the feature database includes: Obtain the D unit characters included in the object aggregation feature, count the occurrence frequency of each unit character in the object aggregation feature, and determine the occurrence frequency corresponding to each unit character as the weight value corresponding to each unit character; D is an integer greater than 1; Construct a feature encoding tree corresponding to the object aggregation feature according to each unit character and the weight value corresponding to each unit character; the leaf nodes in the feature encoding tree are used to represent the D unit characters; the weight value corresponding to the parent node i in the feature encoding tree is the sum of the weight values corresponding to the child nodes of the parent node i; Obtain the shortest path between each leaf node and the root node in the feature encoding tree, and obtain the encoding information of the unit character corresponding to each leaf node according to the position information of each node in the shortest path in the feature encoding tree; Generate an object encoding feature corresponding to the object aggregation feature according to the encoding information corresponding to each unit character and the text position of each unit character in the object aggregation feature, and store the object encoding feature in the feature database.
7. The method according to claim 1, characterized in that The method further includes: Create a data storage class for inheriting the database access interface in the feature generation component, and create a data storage function in the data storage class; Determine the database identification information and data storage structure corresponding to the feature database as the database parameters corresponding to the feature database; Generate the implementation logic of the data storage function according to the database access interface and the database parameters corresponding to the feature database; Execute the implementation logic of the data storage function, and write the database parameters corresponding to the feature database into the service configuration file.
8. The method according to claim 1, characterized in that The method further includes: If a service recommendation request is received, obtain the object identification information corresponding to the service recommendation request; Determine the object aggregation feature in the feature database that matches the object identification information as the object recommendation feature corresponding to the service recommendation request; Obtain the set of services to be recommended corresponding to the service recommendation request, and obtain the service recommendation features corresponding to each service to be recommended in the set of services to be recommended; Obtain the similarity evaluation value between the service recommendation feature corresponding to each service to be recommended and the object recommendation feature, and determine the recommendation status corresponding to each service to be recommended according to the similarity evaluation value; Display the services to be recommended whose recommendation status belongs to the recommendation enabled status in the service application.
9. The method according to claim 8, wherein The number of services to be recommended in the set of services to be recommended is L, and L is a positive integer; The determining the recommendation status corresponding to each service to be recommended according to the similarity evaluation value includes: Sort the similarity evaluation values corresponding to L services to be recommended in descending order to obtain the L services to be recommended after sorting; Obtain the number y corresponding to the recommended display positions in the service application. Among the L services to be recommended after sorting, determine the recommended status of the first y services to be recommended as the recommended enabled status, and determine the recommended status of the services to be recommended other than the first y services to be recommended as the recommended disabled status; y is a positive integer less than or equal to L.
10. A data processing device, characterized in that, The device includes: A database determination module, configured to obtain an initial feature set corresponding to a service application through a feature generation component, and determine a feature database corresponding to the service application according to database parameters included in a service configuration file in the feature generation component; A format conversion module, configured to perform data format conversion on initial object features in the initial feature set to obtain a candidate feature set; the data format corresponding to the candidate object features in the candidate feature set is a data format supported by the feature database; A feature aggregation module, configured to perform feature aggregation processing on the candidate object features in the candidate feature set to obtain an object aggregation feature, and store the object aggregation feature in the feature database; the feature database is used to provide a data source for the service recommendation service corresponding to the service application.
11. A computer device, characterized in that, Including a memory and a processor; The memory is connected to the processor, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, Including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.