Skip to main content

Data Warehousing and ETL Processes

 


1

Goals of Data Warehousing



 1. “The data warehouse must make an organization’s information easily accessible.”


 2. “The data warehouse must present the organization’s information consistently.”


3. “The data warehouse must be adaptive and resilient to change.”


4. “The data warehouse must be a secure bastion that protects out information assets.”


5.“The data warehouse must serve as the foundation for improved decision     making.”


6. “The business community must accept the data warehouse if it is to be deemed successful.”


 

ETL Processes


ETL is an acronym for a three-step process to extract data from source systems and load the data to a data warehouse. The three steps of the ETL process are:



1.Extract and load Staging Tables:


 Extracts and consolidates data from one or more source systems and loads into the data warehouse staging tables.


2.Transforms the data


Transforms data in the staging tables and computes calculated values in preparation for the load.


3.Load Dimension and Fact Tables


Generates and maintains data warehouse surrogate keys and loads target dimension and fact tables.




ETL Processes





ETL processes are further refined to two types of mappings:



 1. Source Dependent Extract (SDE) mappings 


 Extracts the data from the transactional systems and loads to the data warehouse staging tables. SDE mappings are designed with respect to the source’s unique data model.

 



SDE Mapping




2. Source Independent Load (SIL) mappings 


Extracts and transforms data from the staging tables and loads to the data warehouse target tables. SIL mappings are designed to be universal with any source.



SIL Mapping





There is third type mapping known as a Post Load Processing (PLP) mapping, which is used to load aggregate tables after target tables have been loaded. This mapping is designed to be source independent.




ETL Terminology



 ETL – the process by which data is extracted from source A and transformed, aggregated and loaded into source B. ETL can be implemented by such mediums as PL/SQL or various ETL development tools.


 Full Load – the process of extracting all required data from the source and loading it into target tables. A full load will truncate all the target tables.


 Incremental Load – subsequent loads after the full load, extracting source data deltas.


Star Schema – a denormalized schema consisting of a centralized fact and one or more dimensions.


 Snowflake Schema – a normalized schema with multiple facts at different levels of grain and dimensions containing foreign key relationships to other dimensions.

                                                                        11                            

Comments

Popular posts from this blog

Top 100 Informatica Interview Questions

I have attended Informatica interview last week in wipro and couple of other companies, Question below I faced in those companies. 1. What are the main issues while working with flat files as source and as targets ? 2. Explain about Informatica server process that how it works relates to mapping variables? 3. write a query to retrieve the latest records from the table sorted by version(scd) 4. How do you handle two sessions in Informatica 5. which one is better performance wise joiner or look up 6. How to partition the Session? 7. How many types of sessions are there in informatica.please explain them. 8. Explain the pipeline partition with real time example? 9. Explain about cumulative Sum or moving sum? 10. CONVERT MULTIPLE ROWS TO SINGLE ROW (MULTIPLE COLUMNS) IN INFORMATICA 11. DEPLOYMENT GROUPS IN INFORMATICA 12. LOAD LAST N RECORDS OF FILE INTO TARGET TABLE - INFORMATICA 13. LOAD ALTERNATIVE RECORDS / ROWS INTO...

OBIEE 11g dumps

I have cleared OBIEE 11g certification exam (1z0-591) exam yesterday, It was damn easy when compared to OBIEE 10g certification. I have referenced oracle PDF and some of the other OBIEE dumps. I remember some of the question which came in exam. I will publish those question very soon

Configuring the Evaluate Function in OBIEE 12C

/obiee/MW/instances/instance1/config/OracleBIServerComponent/coreapplication_obis1 Change the value of EVALUATE_SUPPORT_LEVEL =0 as below EVALUATE_SUPPORT_LEVEL: 1: evaluate is supported for users with manageRepositories permssion 2: evaluate is supported for any user. other: evaluate is not supported if the value is anything else. EVALUATE_SUPPORT_LEVEL = 2; Configuring Maximum number of records not done for odi server Configuring Maximum number of records Exceeded configured maximum number of allowed input records. Added tags in Instanceconfig.xml file. Steps : 1. login into EM 2.stop the opmn services 3.go to path /obiee/FMW/instances/instance1/config/OracleBIPresentationServicesComponent/coreapplication_obips1 4.edit the instanceconfig.xml file using vi command 5. Add bellow line under views tags <MaxVisibleColumns>30000</MaxVisibleColumns> <MaxVisiblePages>2500000</MaxVisiblePages> <MaxVisibleRows>65000</MaxVisibleRows> <MaxVisib...