Locul din spatele case unde se adună prieteni la o poveste.

Fediverse

Postări de la conturile urmărite

DHH @dhh@bird.makeup
You can see all the default themes here: https://learn.omacom.io/2/the-omarchy-manual/52/themes
Vezi original ↗
Stuart Langridge @sil@mastodon.social
One of the nicest experiences is going to bed while thinking that you have almost all the ingredients for a really great breakfast tomorrow. It’s like being a kid on Christmas Eve.
Vezi original ↗
Manton Reece @manton@manton.org
@lmika I tried today and could get some post text, but no post images and no home page text. Not good. I’m going to test Micro.blog’s archiver to see what I can do there.
Vezi original ↗
Brennan Kenneth Brown @brennan@social.lol
Creating a Digital Garden in 11ty: Tracking My Daily Word Count with URLminder | 🔗 https://brennan.day/creating-a-digital-garden-in-11ty-tracking-my-daily-word-count-with-urlminder/

#Eleventy #Javascript #Technical #Tutorial #IndieWeb #Beeminder #DigitalGarden
Vezi original ↗
DHH @dhh@bird.makeup
Doug can give you a tour. https://www.youtube.com/watch?v=wA3AmChbDU0
Vezi original ↗
Jarkko Sakkinen @jarkko@social.kernel.org
@eduzsh I'm going to keep in phase of doing at min two inferences engines annually from scratch but still polish each to as high quality I can. What I want to understand is this:

1. Let's assume we have a SoC.
2. Let's imagine it is capable of doing inference and has special features.
3. The micro-architecture design can make any feasible sacrifices on anything related training post-training (does not have to but I don't give it any weight).

What would be best architecture provide lift up for let's say up 500B parameter models. It's also definitely an area were Nvidia dose not have any tech leadership. Blackwell hardware design is sloppy and dysoptimal if thinking from this "you had one job" angle.

Next model I'm still going to do on Ryze 5 Pro (common laptop CPU from decade ago, Zen 2 architecture) I need to make it scale to GPT-OSS-120B. That's my end goal for this CPU. I'm planning to reach it with 2bit quantization. I have full MoE implementation for 20B version. The magical "model streaming" part was weird. This was discussed either in the context Dwarf Star 4 or Colibii. I mean one always mmaps huge files instead of copying anything and page fault handler brings up the "experts". Still don't get what model streaming is but I'd guess it is just a silly term for the most common activity (never checked this from their implementation).
Vezi original ↗
Jarkko Sakkinen @jarkko@social.kernel.org
@eduzsh I think lot of problems around AI should be ripped away from AI researchers, we should put them deep into the cellar and let them work only on training at most ;-)

E.g., inference as an algorithm optimization exercise is more like comparable to a driver design than anything to do with machine learning. For me writing couple of inference engines from scratch has been mostly fun and I'm quickly becoming good at it despite I know almost nothing about training and machine learning. However, I think about e.g. CPU cache hierarchy almost every day (very first time was at high school while writing texture mappers for Pentium with its dual-integer pipeline plus additional simultaneous FPU 1/w op at best).

AI researcher can be considred like hardware designer or compiler writer they just have now overemphasized weight. Often e.g., a C++ compiler and successful programs written in C++ come from different carbon based entities :-)
Vezi original ↗

Comenzi rapide de la tastatură

j
Derulează în jos
k
Derulează în sus
d
Derulează jumătate de pagină în jos
u
Derulează jumătate de pagină în sus
gg
Mergi sus de tot
G
Mergi jos de tot
/
Focalizează căutarea
r
Articol aleatoriu
[
Pagina anterioară
]
Ultima pagină
?
Afișează acest ajutor