Use este identificador para citar ou linkar para este item: http://www.repositorio.ufop.br/jspui/handle/123456789/1677
Título: The evolution of web content and search engines.
Autor(es): Yates, Ricardo Baeza
Pereira Junior, Álvaro Rodrigues
Ziviani, Nivio
Data do documento: 2006
Referência: YATES, R. B.; PEREIRA JUNIOR, A. R.; ZIVIANI, N. The evolution of web content and search engines. In. 8th ACM Workshop on Web Mining and Web Usage Analysis,8. 2006. Philadelphia. Anais... Philadelphia: ACM Workshop on Web Mining and Web Usage Analysis, 2006. v. 1. Disponível em: <http://webmining.spd.louisville.edu/webkdd06/papers/paper-7-The%20evolution%20of%20Web-Alvaro-Baeza-Final.pdf>. Acesso em: 18 out. 2012.
Resumo: The evolution of web content and search engines The Web grows at a fast pace and little is known about how new content is generated. The objective of this paper is to study the dynamics of content evolution in the Web, giv-ing answers to questions like: How much new content has evolved from the Web old content? How much of the Web content is biased by ranking algorithms of search engines? We used four snapshots of the Chilean Web containing documents of all the Chilean primary domains, crawled in four distinct periods of time. If a page in a newer snapshot has content of a page in an older snapshot, we say that the source is a parent of the new page. Our hypothesis is that when pages have parents, in a portion of pages there was a query that related the parents and made possible the creation of the new page. Thus, part of the Web content is biased by the ranking function of search engines. We also de¯ne a genealogical tree for the Web, where many pages are new and do not have parents and others have one or more parents. We present the Chilean Web genealogical tree and study its components. To the best of our knowledge this is the ¯rst paper that studies how old content is used to create new content, relating a search engine ranking algorithm with the creation of new pages.
URI: http://www.repositorio.ufop.br/handle/123456789/1677
Aparece nas coleções:DECOM - Trabalhos apresentados em eventos

Arquivos associados a este item:
Arquivo Descrição TamanhoFormato 
EVENTO_EvolutionContentSearch.pdf511,41 kBAdobe PDFVisualizar/Abrir


Os itens no repositório estão protegidos por copyright, com todos os direitos reservados, salvo quando é indicado o contrário.